Please confirm you are human

This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.

A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.

Hold with a pointer, or hold Space or Enter.

News

BenchLM
benchlm.ai > compare > gemini-2-5-pro-vs-gpt-5-5

Gemini 2.5 Pro vs GPT-5.5: Benchmarks, Pricing, Speed (July 2026)

15+ hour, 59+ min ago   (467+ words) Head-to-head evidence from 23 shared benchmark results across 7 categories. Overall scores shown here use the public BenchAlign v5 ranking lane. Public leaderboard positions: Gemini 2.5 Pro #75 (Supported); GPT-5.5 #11 (Estimated). Intervals and evidence labels describe ranking uncertainty, not a guarantee for a specific workload....

BenchLM
benchlm.ai > compare > composer-2-5-vs-ling-3-0-flash

Composer 2.5 vs Ling 3.0 Flash: Benchmarks, Pricing, Speed (July 2026)

15+ hour, 59+ min ago   (309+ words) Head-to-head evidence from 0 shared benchmark results across 0 categories. Overall scores shown here use the public BenchAlign v5 ranking lane. Evidence parity. Composer 2.5 and Ling 3.0 Flash share 0 comparable benchmark results. 0 of 8 categories are comparable. 5 results are unique to Composer 2.5; 0 to Ling 3.0 Flash....

BenchLM
benchlm.ai > compare > command-a-plus-vs-sakana-fugu-ultra

Command A+ vs Sakana Fugu-Ultra: Benchmarks, Pricing, Speed (July 2026)

11+ hour, 59+ min ago   (387+ words) Head-to-head evidence from 1 shared benchmark result across 1 category. Overall scores shown here use the public BenchAlign v5 ranking lane. Public leaderboard positions: Command A+ #142 (Estimated); Sakana Fugu-Ultra unranked (Not scored). Intervals and evidence labels describe ranking uncertainty, not a guarantee for…...

BenchLM
benchlm.ai > compare > gpt-5-6-sol-vs-qwen3-5-122b-a10b

GPT-5.6 Sol vs Qwen3.5-122B-A10B: Benchmarks, Pricing, Speed (July 2026)

15+ hour, 59+ min ago   (563+ words) Head-to-head evidence from 19 shared benchmark results across 6 categories. Overall scores shown here use the public BenchAlign v5 ranking lane. Evidence parity. GPT-5.6 Sol and Qwen3.5-122B-A10B share 19 comparable benchmark results. 5 of 8 categories are comparable. 29 results are unique to GPT-5.6 Sol; 12 to Qwen3.5-122B-A10B. Pick GPT…...

BenchLM
benchlm.ai > compare > gpt-5-6-luna-vs-holo2-4b

GPT-5.6 Luna vs Holo2-4B: Benchmarks, Pricing, Speed (July 2026)

15+ hour, 59+ min ago   (324+ words) Head-to-head evidence from 0 shared benchmark results across 0 categories. Overall scores shown here use the public BenchAlign v5 ranking lane. Public leaderboard positions: GPT-5.6 Luna #24 (Estimated); Holo2-4B unranked (Not scored). Intervals and evidence labels describe ranking uncertainty, not a guarantee for a specific…...

BenchLM
benchlm.ai > benchmarks > aaGlobalMmluLite

AA Global-MMLU-Lite Leaderboard & Scores — July 2026

15+ hour, 59+ min ago   (268+ words) BenchLM mirrors the published score view for AA Global-MMLU-Lite. Gemini 3.1 Pro leads the public snapshot at 93.2%, followed by Gemini 3 Pro (92.2%) and Claude Opus 4.6 (Adaptive) (92.2%). BenchLM does not use these results to rank models overall. Claude Opus 4.6 (Adaptive) The published AA…...

BenchLM
benchlm.ai > benchmarks > aaHarveyLab

AA Harvey LAB Leaderboard & Scores — July 2026

11+ hour, 59+ min ago   (263+ words) BenchLM mirrors the published score view for AA Harvey LAB. Kimi K3 leads the public snapshot at 94.6%, followed by Claude Fable 5 (93.6%) and Muse Spark 1.1 (93.1%). BenchLM does not use these results to rank models overall. The published AA Harvey LAB snapshot places…...

BenchLM
benchlm.ai > compare > gpt-5-6-terra-vs-granite-4-0-350m

GPT-5.6 Terra vs Granite-4.0-350M: Benchmarks, Pricing, Speed (July 2026)

15+ hour, 59+ min ago   (358+ words) Head-to-head evidence from 11 shared benchmark results across 5 categories. Overall scores shown here use the public BenchAlign v5 ranking lane. Evidence parity. GPT-5.6 Terra and Granite-4.0-350M share 11 comparable benchmark results. 0 of 8 categories are comparable. 34 results are unique to GPT-5.6 Terra; 0 to Granite…...

BenchLM
benchlm.ai > compare > gpt-5-6-sol-vs-granite-4-0-350m

GPT-5.6 Sol vs Granite-4.0-350M: Benchmarks, Pricing, Speed (July 2026)

15+ hour, 59+ min ago   (358+ words) Head-to-head evidence from 11 shared benchmark results across 5 categories. Overall scores shown here use the public BenchAlign v5 ranking lane. Evidence parity. GPT-5.6 Sol and Granite-4.0-350M share 11 comparable benchmark results. 0 of 8 categories are comparable. 37 results are unique to GPT-5.6 Sol; 0 to Granite…...

BenchLM
benchlm.ai > compare > gpt-5-6-sol-vs-holo3-1-0-8b

GPT-5.6 Sol vs Holo3.1-0.8B: Benchmarks, Pricing, Speed (July 2026)

15+ hour, 59+ min ago   (372+ words) Head-to-head evidence from 0 shared benchmark results across 0 categories. Overall scores shown here use the public BenchAlign v5 ranking lane. Public leaderboard positions: GPT-5.6 Sol #4 (Supported); Holo3.1-0.8B unranked (Not scored). Intervals and evidence labels describe ranking uncertainty, not a guarantee for a specific…...