Models & Benchmarks
Releases, papers, SOTA benchmarks
Independent · Editorial · Analysis
Independent analysis, tutorials, and commentary on models, research, agents, and policy—written by our desk, sourced to primary links, open for comments.
Five lanes across the AI stack — pick a lane and dive in.
New essays and analysis from models to policy — updated as the field moves.
I explain loop engineering—the shift from hand-prompting coding agents to designing systems that prompt, verify, and iterate on their own—with a step-by-step tutorial, maturity ladder, and honest limits.
Read article →
Tsinghua & BAAI's Brainμ model hits Science, mapping memory-sleep links. I see academic rigor, but commercial viability remains unproven.
China launches a full-stack embodied AI sim using domestic GPUs, advancing the 2025-2026 agenda. Ops take: Local silicon cuts latency but risks supply chain fragility.
I analyze Qwen3.6-27B vs 35B-A3B specs to guide open-source adoption. I think benchmarking methodology remains opaque.
Tashi Zhixang unveiled AWE3.0, a non-simulated, non-VLA embodied model. I read the filing; it claims general-purpose capability without teleoperation or simulation tricks.
OpenAI raises $110B at $730B valuation from Nvidia, Amazon, SoftBank. I read the filings; unit economics matter more than hype.
I see how OpenAI's $110B deal traps creators in a chip-compute loop, prioritizing infrastructure over fair licensing and provenance rights.
I read SenseTime's open-source SenseNova-MARS release, a key multimodal search and reasoning model for the 2025-2026 industry timeline.
Musk opens X's algo, calls it 'terrible' but promises updates. I note this transparency is rare for Musk, signaling a shift in how social platforms handle their core AI engines.
GPT-5.2 Pro independently proved a 45-year number theory conjecture, with Terence Tao confirming no errors found.
I read the UC Berkeley research. It frames a dishwasher robot as an extension topic for the 2025-2026 AI industry cycle.
OpenAI's top reasoning expert leaves after building o3/o1/GPT-4/Codex. This exodus signals deep instability in their core R&D team.
Standout essays and tutorials our desk recommends.
Sign in to comment on stories, reply to other readers, and join the discussion on JustGhostIt.