OBSERVATORY DEGRADED

Semantic review is behind discovery. Evidence and Pulse may be incomplete until the backlog is cleared.

2 material first-party/frontier/open-problem candidate(s) have waited more than 6 hours for semantic review.

0 critical overdue · 2 material overdue
Last semantic import 19 Sept, 04:12

EVIDENCE REGISTER

The evidence archive

A chronological register of published signals. Source class, verification state, evidence quality and relevance remain visible before interpretation.

PUBLISHED RECORDS

222 matching evidence records

151–175 / 222
21 Aug 2026Intelligence

NVIDIA strikes $6B Poolside licensing deal to accelerate open-weight Nemotron models

NVIDIA reportedly agreed to pay about $6 billion to license Poolside AI technology, extend offers to more than 100 Poolside employees, and separately invest about $1 billion in the startup as it builds a stronger U.S. open-weight model stack around Nemotron.

ANNOUNCEDWALL STREET JOURNALPrimary source ↗
Relevance8.2/10evidence 78/100
21 Aug 2026Compute & Infrastructure

NVIDIA invests in Cloverleaf to accelerate powered-land development for AI data centers

NVIDIA made a minority investment in Cloverleaf Infrastructure and entered a strategic partnership to accelerate U.S. AI data-center development. Cloverleaf develops power-ready sites and says it has delivered multiple gigawatt-scale projects; the partnership will use NVIDIA DSX to optimize site, power, cooling and compute infrastructure decisions. Financial terms were not disclosed.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.3/10evidence 95/100
21 Aug 2026Intelligence

NVIDIA AVO completes all 183 ARC-AGI-3 public levels with Claude Opus 5

NVIDIA reports that its Agentic Variation Operators (AVO) long-horizon agent architecture, paired with Claude Opus 5, achieved a 100.00 RHAE score across all 25 ARC-AGI-3 public environments and all 183 levels. The result is system-level, not a controlled model-only uplift: ARC Prize separately reports about 30.2% for Claude Opus 5 at High reasoning effort, while NVIDIA used a different agent system, reasoning setting and text-grid observation setup. No semi-private or fully private ARC-AGI-3 result is reported.

PRIMARY CONFIRMEDNVIDIA TECHNICAL BLOGPrimary source ↗
Relevance8.8/10evidence 90/100
21 Aug 2026Intelligence

OpenAI cuts GPT-5.6 Sol API pricing for a three-month promotional window

OpenAI reduced GPT-5.6 Sol developer pricing from $5 to $4 per million input tokens and from $30 to $20 per million output tokens, a 20% input and 33.3% output reduction, with promotional pricing available at least through November 21, 2026. This lowers the cost of sustained frontier-model workloads but does not change model capability.

INDEPENDENTLY CONFIRMEDAWS / OPENAI ANNOUNCEMENT CONFIRMATIONPrimary source ↗
Relevance7.0/10evidence 92/100
21 Aug 2026Intelligence

Anthropic CHIVE finds no predictive uplift from three activation-reading interpretability tools

Anthropic introduces CHIVE, an agentic counterfactual-evaluation pipeline for explaining naturally occurring LLM behavior. On its checkable proxy evaluation, activation oracles, natural-language autoencoders and sparse autoencoders do not improve prediction of counterfactual behavior over a transcript-only baseline. The authors stress that the proxy differs from harder system-card use cases, so the result is evidence of an interpretability/control bottleneck rather than proof that activation-reading methods are useless.

PRIMARY CONFIRMEDANTHROPIC ALIGNMENT SCIENCEPrimary source ↗
Relevance7.2/10evidence 90/100
21 Aug 2026Intelligence

DeepSeek releases V4-Flash-Vision-Exp with multimodal agent capabilities

DeepSeek released an experimental multimodal version of V4 Flash on its API, adding image understanding while retaining the text model’s reasoning and agent capabilities. DeepSeek reports large gains on vision-dependent agent benchmarks and says multimodal agent performance approaches Claude Opus 4.8; independent benchmark confirmation is still pending.

PRIMARY CONFIRMEDDEEPSEEK API DOCSPrimary source ↗
Relevance7.6/10evidence 90/100
21 Aug 2026Intelligence

Anthropic expands Mythos 5 cyber-defense access through Claude Security

Anthropic made Claude Mythos 5 available for repository vulnerability scans in Claude Security, returning structured findings and suggested fixes without exposing raw model access; it also announced partner integrations, a $35M Defender Advantage Fund, and planned expansion of trusted cyber access.

PRIMARY CONFIRMEDANTHROPICPrimary source ↗
Relevance7.3/10evidence 90/100
20 Aug 2026Compute & Infrastructure

Broadcom seeks $60B+ debt package for frontier-AI chip financing

Broadcom is reported to be in talks with lenders for more than $60B of debt, potentially reaching about $100B across tranches, to finance AI-chip deployments benefiting Anthropic and other frontier labs. The talks extend the $35B Broadcom/Apollo/Blackstone AI XPV platform announced in June.

ANNOUNCEDPRESS / WIREPrimary source ↗
Relevance7.2/10evidence 72/100
20 Aug 2026AI for Science

ChatGPT 5.6 Sol Ultra autonomously supplies central proof idea for first smooth random fast dynamo

Keefer Rowan constructs the first smooth random fast dynamo on T^3. The paper states that the central proof idea was generated autonomously by ChatGPT 5.6 Sol Ultra, while the author wrote and verified the manuscript. The result is a random-flow analogue and does not fully solve the deterministic smooth fast-dynamo problem.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.5/10evidence 84/100
20 Aug 2026Intelligence

Hume AI and Hugging Face quantify benchmark-conditioned behavior in leading ASR models

Researchers introduce three behavioral probes showing that several high-performing open-source speech-recognition models reproduce benchmark-specific reference text even when the audio contradicts, masks, or underdetermines that text. The result indicates that public benchmark scores can overstate general-purpose transcription capability and strengthens the case for held-out, temporally separated evaluations.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.0/10evidence 90/100
20 Aug 2026Intelligence

EnvHarness turns static agent environments into adaptive training layers

Google-affiliated researchers introduce EnvHarness and EnvRigger, a programmable layer that reshapes existing agent environments from observed failure trajectories while preserving the original verifier. Across five benchmarks in four domains, the paper reports up to a 9.0 percentage-point improvement on held-out instances and 9.8% fewer execution steps.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.3/10evidence 86/100
20 Aug 2026AI for Science

AI-assisted counterexample shows Marton inner bound is strictly sub-optimal

A 20 Aug 2026 preprint by Mian Huang, Yanxiao Liu and Yi Liu gives an unconditional counterexample showing the complete one-letter Marton region can be strictly sub-optimal. The numerical search relied heavily on GPT-5.6 Sol, Claude Fable 5 and Opus 5; the authors certify the numerical gap with exact rational arithmetic and outward-rounded MPFR.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.4/10evidence 86/100
20 Aug 2026Intelligence

Thinkingbox exposes a large reliability gap in stateful business agents

A new Microsoft-linked arXiv benchmark evaluates 507 stateful, policy-conditioned business workflows with executable checks on final backend state and side effects. The strongest tested model reaches 65.36% pass@1 but succeeds in all 20 repeated attempts on only 25.25% of tasks, showing that occasional agent success substantially overstates dependable execution.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.6/10evidence 92/100
20 Aug 2026Intelligence

Ox Alpha stealth model appears on OpenRouter with 1M context and multimodal agentic focus

OpenRouter released an anonymous stealth preview called Ox Alpha on August 20. Its official listing confirms a 1,048,576-token context window, text/image/video input, tool use, structured outputs and positioning for coding and sustained agentic work. Community claims that it beats leading frontier models are not independently validated, and the developer identity remains undisclosed.

PRIMARY CONFIRMEDOPENROUTERPrimary source ↗
Relevance7.4/10evidence 88/100
20 Aug 2026Intelligence

Meta details Muse Spark 1.2 multimodal, agentic and robotics capabilities

Meta published new Muse Spark 1.2 evaluations and demonstrations spanning visual coding, audio-visual understanding, tool-augmented multimodal reasoning and a specialized robotics stack. The robotics setup uses Muse Spark as a high-level planner with a lower-level VLA policy for manipulation; Meta also previewed WildArtifactBench for real-world multimodal agent utility.

PRIMARY CONFIRMEDMETA AI RESEARCHPrimary source ↗
Relevance7.4/10evidence 88/100
20 Aug 2026AI for Science

Elliptic-curve rank record reaches at least 30

A newly submitted curve on the iCARM Elliptic Curve Rank Leaderboard has rank at least 30, satisfying a named AIM problem that asked for an elliptic curve over Q of rank 30. The leaderboard uses exact 2-descent certification to establish independence of the exhibited rational points. Public discussion links the discovery to Levent Alpöge, Ava Howell and Claude, but the discovery method and AI contribution have not yet been documented in a primary paper, so the AI role remains unclear.

PRIMARY CONFIRMEDICARM ELLIPTIC CURVE RANK LEADERBOARDPrimary source ↗
Relevance8.8/10evidence 88/100
20 Aug 2026Intelligence

Anthropic makes computer use, browser use, Skills API and Files API production-ready

Anthropic moved computer use, the Skills API and the Files API to general availability and added a new browser-use tool that combines screenshots with page structure. Computer use can now execute several actions per turn; Anthropic reports a customer claims workflow falling from 32 to 13 minutes with 100% completion in its tests.

PRIMARY CONFIRMEDANTHROPICPrimary source ↗
Relevance7.1/10evidence 88/100
19 Aug 2026AI for Science

Claude-led algorithm breaks the 2^n barrier for exact linear-extension counting

Keigo Oka gives a deterministic exact O*(1.89^n) algorithm for counting linear extensions of arbitrary n-element posets, resolving Koivisto’s 2013 Dagstuhl question. The paper disclosure says Claude Opus 5 discovered the core mathematical ideas behind the new algorithm and proof; the research prompt was generated by ChatGPT 5.6 Sol.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.3/10evidence 86/100
19 Aug 2026AI for Science

GPT-5.6 Sol Pro supplies proof resolving first open Big-Line-Big-Clique case

Édouard Bonnet proves that every sufficiently large finite planar point set has either four collinear points or six pairwise visible points, resolving the first open case of the Kára–Pór–Wood Big-Line-Big-Clique conjecture. The paper states that after a failed attempt and generic encouragement, GPT-5.6 Sol Pro produced a relatively detailed proof after 222 minutes; the author checked and simplified it and rewrote the exposition.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.6/10evidence 86/100
19 Aug 2026Intelligence

Guidelight assessment finds major gaps in frontier-AI control and containment

Guidelight AI Standards assessed public safety and control practices at OpenAI, Anthropic, Google, xAI and Meta; Reuters reports the best grades were only C+ and highlights insufficient preventive controls and containment across the sector. This is a governance/control bottleneck signal, not a new capability result.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.2/10evidence 84/100
19 Aug 2026AI for Science

AI-assisted proof completes the Gaussian boson sampling hiding conjecture

Shou, Gorshkov, Galitski and Miller report a complete proof of the hiding conjecture for Gaussian boson sampling with an arbitrary number of squeezed input modes. The acknowledgments explicitly state that GPT-5.5 Thinking/Pro and GPT-5.6 Sol were used to generate proof ideas and methods, with all results checked and validated by the human authors.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.7/10evidence 88/100
19 Aug 2026Compute & Infrastructure

Google expands custom-AI silicon partnership with Marvell

Marvell will help develop Google custom AI chips and granted Google warrants for up to 58.97 million shares worth about $12.2B if fully exercised. The arrangement could generate up to roughly $120B of Marvell revenue through fiscal 2033 if order targets are met, broadening Google TPU supply across processors, storage and networking.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.6/10evidence 95/100
19 Aug 2026Synthetic Biology

Intismeran personalized mRNA therapy meets Phase 3 melanoma endpoints

Moderna and Merck reported that individualized mRNA neoantigen therapy intismeran autogene plus pembrolizumab met the primary recurrence-free-survival endpoint and the key distant-metastasis-free-survival endpoint in the 1,137-patient Phase 3 INTerpath-001 study after resection of high-risk stage IIB-IV melanoma. Full effect sizes and detailed data have not yet been released.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance8.2/10evidence 88/100
19 Aug 2026Robotics

Unitree completes Shanghai listing in a major robotics capital-scale milestone

Unitree began trading on Shanghai's STAR Market on 19 August 2026 after raising about $900 million. Shares closed 460% above the IPO price, valuing the company at roughly $50 billion. This is a capital and industrial-scaling signal, not evidence of a new autonomy or embodied-intelligence capability.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.0/10evidence 92/100
19 Aug 2026AI for Science

Eureka meta-agent reports long-horizon scientific-discovery gains

ManXis reports a task-conditioned meta-agent architecture that dynamically forms specialized scientific agents, completing 170/170 recursive long-horizon tasks and producing certified mathematical/theoretical outputs. The paper also reports progress on a localized Weil-positivity certificate related to the Riemann Hypothesis, explicitly stating it is not an RH proof.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.6/10evidence 76/100