Model intelligence / Live profile

Claude

Track Claude from Anthropic across launches, availability, pricing, benchmarks, capabilities, and consequential updates. This page updates automatically as substantive Superpower Daily coverage is published.

Official page
Coverage telemetryActive

Developer / company

Anthropic
Published28
Launches6
Pricing0
Benchmarks4
Published coverage28 stories
Last 90 days21 updates
Coverage dates18 days
Profile statusIndexed

Latest coverage

What changed around Claude.

Browse all models

Update timeline

The maintained record.

BenchmarkStanford Won Databricks’ Agent Cup, but 18.8% of Questions Stumped Every TeamThe live contest tested whether agents developed on one benchmark could handle an unfamiliar Treasury archive. Stanford led the field, but the unsolved questions show the limits of reliable document reasoning.LaunchDatabricks Says Its New Extraction Mode Beats Frontier Models on the Documents That Break ThemPrecision Mode is designed for cross-page reasoning, enormous line-item outputs and deeply nested schemas. Databricks reports a seven-point lead over its best tested chunk-and-merge baseline, but the evaluation blends internal and public datasets.BenchmarkGLM-5.3’s Cheap Retries Put Fable 5’s Coding Premium Under PressureTogether AI’s benchmark run suggests first-attempt parity is not enough to justify Fable 5’s price for general coding. But a Rust and serialization advantage, plus benchmark-specific cost assumptions, leave room for targeted routing.Platform shiftAWS’s RAG Cost Cut Comes With a 19% Latency BillAWS’s two-model pattern sharply reduced the context reaching its answer model in one benchmark. The tradeoff is another inference step, a modest quality decline, and results that may not transfer to a company’s own documents.BenchmarkAI21’s 8B Verifier Challenges the Case for Bigger Search ModelsAI21’s company-reported tests suggest an independently trained checking model can recover answers an ensemble already found but failed to choose. The results are promising, but they rest on 100-question samples, automated grading, and a verifier trained partly through closed-model distillation.BenchmarkAgentX Replays Claude Code Sessions to Test AI Serving SystemsThe replay dataset is designed to measure infrastructure behind agentic coding workloads, not whether a model writes better code. Its value depends on whether synthetic traces preserve the traffic patterns operators need to serve.RegulationClaude’s New Text Watermark Will Steer Word Choices—and Test Whether Anyone NoticesAnthropic says its cryptographic text watermark will not affect quality, speed, or cost. The method’s limits—short passages, constrained answers, and substantial edits—will determine whether it becomes useful provenance infrastructure or a fragile compliance layer.RegulationAnthropic Says Claude’s Planned Watermark May Survive Light Edits—but Not a Complete RewriteA planned API will check a pattern encoded through subtle word choices, but constrained code, lightly edited human drafts and complete rewrites expose the boundaries of what the signal can establish.Platform shiftChatGPT brings unlimited text chats to free usersSan Francisco artist’s fake AI chatbot goes viral — because it’s actually humanSecurity riskClaude Helped a Hacker Find a Way to Issue Tickets to Almost Every US Music FestivalAnthropic wants to develop its own drugs

Your watchlist

Follow every consequential Claude update.