OBSERVATORY DEGRADED

Semantic review is behind discovery. Evidence and Pulse may be incomplete until the backlog is cleared.

2 material first-party/frontier/open-problem candidate(s) have waited more than 6 hours for semantic review.

0 critical overdue · 2 material overdue
Last semantic import 19 Sept, 04:12

EVIDENCE REGISTER

The evidence archive

A chronological register of published signals. Source class, verification state, evidence quality and relevance remain visible before interpretation.

PUBLISHED RECORDS

215 matching evidence records

176–200 / 215
19 Aug 2026Synthetic Biology

FDA approves first therapy for glycogen storage disease type Ia

FDA granted accelerated approval to Genglycos (pariglasgene brecaparvovec-opnr), a one-time AAV8 gene therapy delivering a functional G6PC gene to the liver for patients age 8+ with GSDIa. In a randomized placebo-controlled study, treated patients had a statistically significant 31% mean reduction in daily cornstarch intake; confirmatory trials are still required.

PRIMARY CONFIRMEDFDAPrimary source ↗
Relevance7.4/10evidence 95/100
19 Aug 2026Robotics

Generalist GEN-1.5 demonstrates one-shot physical task learning

Generalist AI reports that GEN-1.5 can execute previously unseen short-horizon manipulation tasks after a single 3–12 second physical demonstration with no gradient update, averaging 59% success across 10 internal tasks; few-step adaptation reaches 83% after about five minutes of task data. The company also reports compositional prompting, sim-to-real prompting and some human-to-robot transfer. Results are first-party and success remains modest on short atomic tasks.

PRIMARY CONFIRMEDGENERALIST AIPrimary source ↗
Relevance8.6/10evidence 82/100
19 Aug 2026AI for Science

Autonomous AI scientist reaches 92.2% on Deep Origin DO Challenge

Deep Origin reran its computational drug-discovery challenge with 2026 frontier models. The best autonomous run recovered 922/1000 hidden top structures, above the prior unrestricted human reference of 77.8%; five runs exceeded that reference. Performance is strongly harness-dependent and the public benchmark may have been exposed in training, limiting causal inference.

PRIMARY CONFIRMEDDEEP ORIGINPrimary source ↗
Relevance8.3/10evidence 88/100
19 Aug 2026AI for Science

Erdős Problem #501 resolved as independent of ZFC with Lean-verified human–AI work

A Lean 4 development resolves the first question of Erdős Problem #501 as independent of ZFC: CH gives a negative model, while adding sufficiently many random reals gives a positive model. The missing large-cardinal-free transfer is credited to Elliot Glazer working with Sol and Claude; prior components were human results.

PRIMARY CONFIRMEDFORMAL / REPOSITORYPrimary source ↗
Relevance7.3/10evidence 92/100
18 Aug 2026Compute & Infrastructure

Marin launches a fully open 535B-A23B frontier-scale training run

The Marin community launched a 535.3B-total / 22.76B-active MoE training run with an 18T-token budget on 11 GB200 NVL72 racks, publishing the model configuration, scaling ladder, training telemetry and research process in public. The signal is transparency and open frontier-scale training infrastructure, not a demonstrated capability jump.

PRIMARY CONFIRMEDFORMAL / REPOSITORYPrimary source ↗
Relevance7.1/10evidence 95/100
18 Aug 2026Intelligence

OpenAI broadens Astra safety slowdown and rewrites frontier testing controls

OpenAI disclosed a broader operational slowdown around frontier-model work: two weeks of deployment-focused RL/model testing were paused, its largest planned frontier RL run remains on hold, stronger sandboxing and AI-based monitoring are being added, and the Preparedness Framework is being rewritten as models approach its critical cyber thresholds.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.9/10evidence 92/100
18 Aug 2026AI for Science

Stein dimension-free weak-(1,1) Riesz transform problem resolved

Ouyang, Spector and Stockdale prove a dimension-free weak-type (1,1) bound with constant 2 for the vector Riesz transform, settling Stein’s 1986 ICM problem. The paper states that the proof strategy was developed through LLM dialogues involving GPT-5.6 Sol, OpenAI reasoning agents and Claude Opus 5.0; the authors then independently checked, rewrote and validated the argument.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.4/10evidence 82/100
18 Aug 2026AI for Science

GPT-5.6 Sol-assisted work disproves Sato's weak F-equivalence conjecture

Avik Chakravarty, Daebeom Choi and Shengjing Xu construct counterexamples disproving Sato's weak F-equivalence conjecture for nonsingular projective toric weak Fano varieties in every dimension at least three, then formulate a Gorenstein refinement. The paper states that the results were developed with assistance from GPT-5.6 Sol.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.7/10evidence 88/100
18 Aug 2026AI for Science

Human–AI collaboration produces new representation-theory results with full Lean formalization

Haruhisa Enomoto reports a research collaboration with Fable 5 and GPT-5.6 Sol that proposes a quotient–submodule equidistribution conjecture and proves multiple nontrivial families, with a full Lean 4 formalization. This is evidence of frontier models participating in novel mathematical research, but it does not close a pre-existing named open problem.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.8/10evidence 91/100
18 Aug 2026Synthetic Biology

GenBio AI launches AIDO Cell general-purpose virtual-cell simulator

GenBio AI introduced AIDO Cell, a multiscale simulator integrating multimodal biological data into a shared world model intended to simulate cellular responses and support therapeutic design. The company reports literature-based validation and says prospective wet-lab validation of novel predictions is underway.

PRIMARY CONFIRMEDGENBIO AIPrimary source ↗
Relevance7.8/10evidence 72/100
18 Aug 2026AI for Science

Claude autonomously designs de novo protein binders validated in wet lab

Anthropic reports that Claude autonomously orchestrated open-source protein-design and folding tools from a human-written research prompt and produced binders against 14 of 15 targets. Adaptyv Bio and Twist Bioscience independently synthesized and tested the designs; reported hit rates were 22–35%, above the 10–15% field baseline cited by Anthropic.

INDEPENDENTLY CONFIRMEDFIRST-PARTY SOCIALPrimary source ↗
Relevance8.8/10evidence 88/100
18 Aug 2026Compute & Infrastructure

Cerebras launches CS-4 wafer-scale inference system for hyperscale AI

Cerebras introduced CS-4, a rack-scale system built from three WSE-3 Turbo processors. The company reports up to 30x faster inference than production GPU systems and up to 10x more throughput per watt than CS-3; first shipments are scheduled for Q3 2026. The strongest headline figures are partly based on internal benchmarking and projections, so they should not be treated as independently validated capability measurements.

PRIMARY CONFIRMEDCEREBRASPrimary source ↗
Relevance7.3/10evidence 86/100
17 Aug 2026AI for Science

FAR pipeline scales AI-assisted mathematical discovery across thousands of open problems

The Find-Attempt-Recommend pipeline scans thousands of combinatorics papers, identifies thousands of apparently open conjectures, surfaces hundreds of potential resolutions, and narrows them to a small expert-review set. This is a material AI-for-Science signal because it automates problem selection and triage, two human bottlenecks in research-level mathematics.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.6/10evidence 82/100
17 Aug 2026Compute & Infrastructure

NVIDIA backs OpenAI Ohio AI campus with up to $105B in lease guarantees

NVIDIA agreed to guarantee up to $105B in lease payments for an OpenAI data-center campus in Pike County, Ohio, and to invest $1.5B in SB Energy. The initial 4.25 GW phase is expected to come online from 2028, with an option for total capacity to reach 8 GW. This is a compute/power/capital scaling signal, not a new model-capability result.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.8/10evidence 90/100
16 Aug 2026AI for Science

Talagrand convolution conjecture solved with AI-discovered proof

A new preprint claims a full proof of Talagrand's 1989 convolution conjecture on the Boolean hypercube. The authors state that the proof was discovered by the Odin Automatic AI Research Agent; they reorganized the final proofs. Independent mathematical review is still pending.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.1/10evidence 82/100
14 Aug 2026Intelligence

Z.ai announces GLM-5.3 with strong cyber-defense results ahead of public release

Z.ai says its upcoming open-weight GLM-5.3 reached 84.5% on CyberGym and will be released publicly in about two weeks after additional security work; the benchmark claims are not yet independently verified.

PRIMARY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.2/10evidence 88/100
14 Aug 2026Robotics

Unitree’s heavily oversubscribed IPO becomes a robotics scale-and-capital signal

Unitree’s Shanghai STAR Market IPO raised about 6.1 billion yuan while the retail tranche was reported more than 8,000 times oversubscribed. Independent reporting also describes Unitree as profitable and among the largest humanoid suppliers by sales. This is evidence that general-purpose robotics can attract substantial capital and manufacturing scale, but it is not evidence of improved autonomy, reliability, intervention rate or generalized task performance.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.0/10evidence 82/100
14 Aug 2026Intelligence

Anthropic reports significant internal AI-R&D speedups but says automated-R&D threshold is not met

Anthropic’s August 2026 first-party Risk Report says its current models are providing significant speedups to AI research and engineering, while explicitly concluding that they do not yet fully substitute for Anthropic researchers and have not clearly doubled the overall pace of progress. The report also names internally deployed Mythos-class systems used in the assessment.

PRIMARY CONFIRMEDANTHROPICPrimary source ↗
Relevance8.4/10evidence 96/100
13 Aug 2026Intelligence

Google launches Gemini 3.7 Flash for coding and agent workflows

Google released Gemini 3.7 Flash on 13 August 2026 for software coding, multi-step agentic workflows and automated business tasks. Independent reporting describes meaningful gains in planning, instruction following and code generation, with lower pricing than the preceding Flash release.

INDEPENDENTLY CONFIRMEDPRESS / WIREPrimary source ↗
Relevance7.5/10evidence 80/100
13 Aug 2026AI for ScienceBreakthrough

ChatGPT 5.6 Sol-assisted construction resolves elliptic regularity question negatively

Nam Q. Le, Qi Sun and Hung V. Tran construct uniformly elliptic nondivergence-form equations in dimension three whose bounded solutions have unbounded W1,1 variation, showing that no interior W1,p estimate depending only on ellipticity exists for p at least one and resolving negatively an open question raised by Nadirashvili, Tkachev and Vlăduţ. The authors state that the main results were obtained through chats with ChatGPT 5.6 Sol and that key strategies came from ChatGPT; they then reworked, rewrote and checked all arguments. The result remains without independent external proof reproduction.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.3/10evidence 89/100
13 Aug 2026AI for Science

Inherent Faraday outperforms frontier agents on a paper-replication benchmark

Inherent reports that Faraday, a 27B AI Scientist agent trained with long-horizon RL and using coding agents as tools, outperforms Claude Opus 4.8 and GPT-5.5 on held-out scientific paper-replication tasks in the new 310-task Replica benchmark. The result is first-party and benchmark/judge design is controlled by the same team, so independent replication remains pending.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.6/10evidence 84/100
13 Aug 2026Intelligence

DeepSeek V4 Pro reaches GA with a major agent-capability upgrade

DeepSeek released the GA version of V4 Pro on 13 August 2026 across app, web and API. DeepSeek reports substantial agent gains, including Terminal Bench 2.1 at 87.9, HLE at 42.7 without tools and 60.0 with tools, Toolathlon-Verified at 74.1 and DSBench-FullStack at 71.1. The API also adds native Responses API support and low/high/max thinking effort. This is a material frontier-model release, but independent evaluation is still needed before treating it as a new global capability ceiling.

PRIMARY CONFIRMEDDEEPSEEK API DOCSPrimary source ↗
Relevance8.0/10evidence 88/100
12 Aug 2026Intelligence

Grok 4.6 launches with a material jump in frontier and agentic performance

SpaceXAI released Grok 4.6 on 12 August 2026. Independent reporting citing Artificial Analysis says the model improves roughly five points over Grok 4.5 on the Artificial Analysis Intelligence Index, lands around GPT-5.6 Sol overall, and is especially strong on agentic work while remaining materially cheaper than the top Anthropic models. This strengthens SpaceXAI's frontier position, but it does not currently establish a new absolute capability ceiling.

INDEPENDENTLY CONFIRMEDSPACEXAI / ELON MUSK; CROSS-CHECKED VIA MARKETWATCH AND INVESTOR'S BUSINESS DAILYPrimary source ↗
Relevance8.5/10evidence 82/100
11 Aug 2026AI for Science

Long-horizon AI research system helps set new bounds on the Grothendieck constant

A human-AI research collaboration tightened both lower and upper bounds on the Grothendieck constant. The companion methodology paper states that the AI system discovered and first proved the new lower bound K_G >= 6π/11, later independently checked and revised by the human authors; the exact value of K_G remains open.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.4/10evidence 92/100
11 Aug 2026AI for ScienceBreakthrough

Claude-assisted Riemann-related result improves a critical-line bound to 67.2%

A Claude-assisted research workflow reportedly improved the proportion of non-trivial simple zeros shown to lie on the critical line from about 41.6% to 67.2%. This was a substantial partial result, not a proof of the Riemann Hypothesis. The work used roughly 31M output tokens, about 60 sub-agents, around 650 discarded strategies, thousands of computational checks and Lean formalization.

PRIMARY CONFIRMEDANTHROPIC RESEARCH / VIBEMATHED
Relevance9.0/10