OBSERVATORY DEGRADED

Semantic review is behind discovery. Evidence and Pulse may be incomplete until the backlog is cleared.

2 material first-party/frontier/open-problem candidate(s) have waited more than 6 hours for semantic review.

0 critical overdue · 2 material overdue
Last semantic import 19 Sept, 04:12

EVIDENCE REGISTER

The evidence archive

A chronological register of published signals. Source class, verification state, evidence quality and relevance remain visible before interpretation.

PUBLISHED RECORDS

70 matching evidence records

51–70 / 70
21 Aug 2026Intelligence

NVIDIA AVO completes all 183 ARC-AGI-3 public levels with Claude Opus 5

NVIDIA reports that its Agentic Variation Operators (AVO) long-horizon agent architecture, paired with Claude Opus 5, achieved a 100.00 RHAE score across all 25 ARC-AGI-3 public environments and all 183 levels. The result is system-level, not a controlled model-only uplift: ARC Prize separately reports about 30.2% for Claude Opus 5 at High reasoning effort, while NVIDIA used a different agent system, reasoning setting and text-grid observation setup. No semi-private or fully private ARC-AGI-3 result is reported.

PRIMARY CONFIRMEDNVIDIA TECHNICAL BLOGPrimary source ↗
Relevance8.8/10evidence 90/100
21 Aug 2026Intelligence

OpenAI cuts GPT-5.6 Sol API pricing for a three-month promotional window

OpenAI reduced GPT-5.6 Sol developer pricing from $5 to $4 per million input tokens and from $30 to $20 per million output tokens, a 20% input and 33.3% output reduction, with promotional pricing available at least through November 21, 2026. This lowers the cost of sustained frontier-model workloads but does not change model capability.

INDEPENDENTLY CONFIRMEDAWS / OPENAI ANNOUNCEMENT CONFIRMATIONPrimary source ↗
Relevance7.0/10evidence 92/100
21 Aug 2026Intelligence

Anthropic CHIVE finds no predictive uplift from three activation-reading interpretability tools

Anthropic introduces CHIVE, an agentic counterfactual-evaluation pipeline for explaining naturally occurring LLM behavior. On its checkable proxy evaluation, activation oracles, natural-language autoencoders and sparse autoencoders do not improve prediction of counterfactual behavior over a transcript-only baseline. The authors stress that the proxy differs from harder system-card use cases, so the result is evidence of an interpretability/control bottleneck rather than proof that activation-reading methods are useless.

PRIMARY CONFIRMEDANTHROPIC ALIGNMENT SCIENCEPrimary source ↗
Relevance7.2/10evidence 90/100
21 Aug 2026Intelligence

DeepSeek releases V4-Flash-Vision-Exp with multimodal agent capabilities

DeepSeek released an experimental multimodal version of V4 Flash on its API, adding image understanding while retaining the text model’s reasoning and agent capabilities. DeepSeek reports large gains on vision-dependent agent benchmarks and says multimodal agent performance approaches Claude Opus 4.8; independent benchmark confirmation is still pending.

PRIMARY CONFIRMEDDEEPSEEK API DOCSPrimary source ↗
Relevance7.6/10evidence 90/100
21 Aug 2026Intelligence

Anthropic expands Mythos 5 cyber-defense access through Claude Security

Anthropic made Claude Mythos 5 available for repository vulnerability scans in Claude Security, returning structured findings and suggested fixes without exposing raw model access; it also announced partner integrations, a $35M Defender Advantage Fund, and planned expansion of trusted cyber access.

PRIMARY CONFIRMEDANTHROPICPrimary source ↗
Relevance7.3/10evidence 90/100
20 Aug 2026Intelligence

Hume AI and Hugging Face quantify benchmark-conditioned behavior in leading ASR models

Researchers introduce three behavioral probes showing that several high-performing open-source speech-recognition models reproduce benchmark-specific reference text even when the audio contradicts, masks, or underdetermines that text. The result indicates that public benchmark scores can overstate general-purpose transcription capability and strengthens the case for held-out, temporally separated evaluations.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.0/10evidence 90/100
20 Aug 2026Intelligence

EnvHarness turns static agent environments into adaptive training layers

Google-affiliated researchers introduce EnvHarness and EnvRigger, a programmable layer that reshapes existing agent environments from observed failure trajectories while preserving the original verifier. Across five benchmarks in four domains, the paper reports up to a 9.0 percentage-point improvement on held-out instances and 9.8% fewer execution steps.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.3/10evidence 86/100
20 Aug 2026Intelligence

Thinkingbox exposes a large reliability gap in stateful business agents

A new Microsoft-linked arXiv benchmark evaluates 507 stateful, policy-conditioned business workflows with executable checks on final backend state and side effects. The strongest tested model reaches 65.36% pass@1 but succeeds in all 20 repeated attempts on only 25.25% of tasks, showing that occasional agent success substantially overstates dependable execution.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.6/10evidence 92/100
20 Aug 2026Intelligence

Ox Alpha stealth model appears on OpenRouter with 1M context and multimodal agentic focus

OpenRouter released an anonymous stealth preview called Ox Alpha on August 20. Its official listing confirms a 1,048,576-token context window, text/image/video input, tool use, structured outputs and positioning for coding and sustained agentic work. Community claims that it beats leading frontier models are not independently validated, and the developer identity remains undisclosed.

PRIMARY CONFIRMEDOPENROUTERPrimary source ↗
Relevance7.4/10evidence 88/100
20 Aug 2026Intelligence

Meta details Muse Spark 1.2 multimodal, agentic and robotics capabilities

Meta published new Muse Spark 1.2 evaluations and demonstrations spanning visual coding, audio-visual understanding, tool-augmented multimodal reasoning and a specialized robotics stack. The robotics setup uses Muse Spark as a high-level planner with a lower-level VLA policy for manipulation; Meta also previewed WildArtifactBench for real-world multimodal agent utility.

PRIMARY CONFIRMEDMETA AI RESEARCHPrimary source ↗
Relevance7.4/10evidence 88/100
20 Aug 2026Intelligence

Anthropic makes computer use, browser use, Skills API and Files API production-ready

Anthropic moved computer use, the Skills API and the Files API to general availability and added a new browser-use tool that combines screenshots with page structure. Computer use can now execute several actions per turn; Anthropic reports a customer claims workflow falling from 32 to 13 minutes with 100% completion in its tests.

PRIMARY CONFIRMEDANTHROPICPrimary source ↗
Relevance7.1/10evidence 88/100
19 Aug 2026Intelligence

Guidelight assessment finds major gaps in frontier-AI control and containment

Guidelight AI Standards assessed public safety and control practices at OpenAI, Anthropic, Google, xAI and Meta; Reuters reports the best grades were only C+ and highlights insufficient preventive controls and containment across the sector. This is a governance/control bottleneck signal, not a new capability result.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.2/10evidence 84/100
18 Aug 2026Intelligence

OpenAI broadens Astra safety slowdown and rewrites frontier testing controls

OpenAI disclosed a broader operational slowdown around frontier-model work: two weeks of deployment-focused RL/model testing were paused, its largest planned frontier RL run remains on hold, stronger sandboxing and AI-based monitoring are being added, and the Preparedness Framework is being rewritten as models approach its critical cyber thresholds.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.9/10evidence 92/100
14 Aug 2026Intelligence

Z.ai announces GLM-5.3 with strong cyber-defense results ahead of public release

Z.ai says its upcoming open-weight GLM-5.3 reached 84.5% on CyberGym and will be released publicly in about two weeks after additional security work; the benchmark claims are not yet independently verified.

PRIMARY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.2/10evidence 88/100
14 Aug 2026Intelligence

Anthropic reports significant internal AI-R&D speedups but says automated-R&D threshold is not met

Anthropic’s August 2026 first-party Risk Report says its current models are providing significant speedups to AI research and engineering, while explicitly concluding that they do not yet fully substitute for Anthropic researchers and have not clearly doubled the overall pace of progress. The report also names internally deployed Mythos-class systems used in the assessment.

PRIMARY CONFIRMEDANTHROPICPrimary source ↗
Relevance8.4/10evidence 96/100
13 Aug 2026Intelligence

Google launches Gemini 3.7 Flash for coding and agent workflows

Google released Gemini 3.7 Flash on 13 August 2026 for software coding, multi-step agentic workflows and automated business tasks. Independent reporting describes meaningful gains in planning, instruction following and code generation, with lower pricing than the preceding Flash release.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.5/10evidence 80/100
13 Aug 2026Intelligence

DeepSeek V4 Pro reaches GA with a major agent-capability upgrade

DeepSeek released the GA version of V4 Pro on 13 August 2026 across app, web and API. DeepSeek reports substantial agent gains, including Terminal Bench 2.1 at 87.9, HLE at 42.7 without tools and 60.0 with tools, Toolathlon-Verified at 74.1 and DSBench-FullStack at 71.1. The API also adds native Responses API support and low/high/max thinking effort. This is a material frontier-model release, but independent evaluation is still needed before treating it as a new global capability ceiling.

PRIMARY CONFIRMEDDEEPSEEK API DOCSPrimary source ↗
Relevance8.0/10evidence 88/100
12 Aug 2026Intelligence

Grok 4.6 launches with a material jump in frontier and agentic performance

SpaceXAI released Grok 4.6 on 12 August 2026. Independent reporting citing Artificial Analysis says the model improves roughly five points over Grok 4.5 on the Artificial Analysis Intelligence Index, lands around GPT-5.6 Sol overall, and is especially strong on agentic work while remaining materially cheaper than the top Anthropic models. This strengthens SpaceXAI's frontier position, but it does not currently establish a new absolute capability ceiling.

INDEPENDENTLY CONFIRMEDSPACEXAI / ELON MUSK; CROSS-CHECKED VIA MARKETWATCH AND INVESTOR'S BUSINESS DAILYPrimary source ↗
Relevance8.5/10evidence 82/100
10 Aug 2026Intelligence

BDH-CQ sets a new ARC-AGI-1 cost-efficiency point with recurrent latent reasoning

Pathway reports that BDH-CQ, a proprietary 150-million-parameter post-Transformer system, combines in-context learning with recurrent latent reasoning and reaches 29.5% pass@2 on the public ARC-AGI-1 evaluation set. The reported operating point uses about 0.85 H200 GPU-seconds per task, corresponding to a computed inference cost of $0.00070 per task at the paper's hardware-price assumption. A documented black-box audit by external-affiliation co-authors reproduced the deployed system's 29.5% score without access to model weights.

INDEPENDENTLY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.2/10evidence 78/100
04 Aug 2026Intelligence

Qwen3.8-Max highlighted in the Daily Pulse

The 4 August Pulse recorded Qwen3.8-Max at 2.4T parameters and 1M context, relevance 7/10, with no ASI forecast change.

INDEPENDENTLY CONFIRMEDSURGE AI INDEPENDENT EVALUATIONPrimary source ↗
Relevance7.4/10evidence 88/100
Page 3 / 3
← PreviousNext →