OBSERVATORY DEGRADED

Semantic review is behind discovery. Evidence and Pulse may be incomplete until the backlog is cleared.

2 material first-party/frontier/open-problem candidate(s) have waited more than 6 hours for semantic review.

0 critical overdue · 2 material overdue
Last semantic import 19 Sept, 04:12

EVIDENCE REGISTER

The evidence archive

A chronological register of published signals. Source class, verification state, evidence quality and relevance remain visible before interpretation.

PUBLISHED RECORDS

96 matching evidence records

26–50 / 96
07 Sept 2026AI for ScienceBreakthrough

ChatGPT 5.6 Sol generates main proof content for sharp curvature-sign rigidity results

Minbo Gao, Yuhang Liu and Genyuan Zhang establish curvature-sign rigidity and sharp pointwise pinching thresholds for sectional and Ricci curvature, including the sharp Ricci threshold 1/(n-1) and matching subcritical examples. The abstract says the main content of the proof was generated by ChatGPT 5.6 Sol and verified by the authors. The results are primary-source confirmed but not independently externally reproduced or peer reviewed.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.0/10evidence 88/100
07 Sept 2026AI for ScienceBreakthrough

GPT-6 Astra generates most arguments in result on AMP and low-degree polynomial equivalence

Zhangsong Li proves an almost-sharp equivalence between approximate message passing and growing-degree polynomial estimation for the Bernoulli rank-one planted-submatrix setting, resolving the Bernoulli rank-one case of a growing-degree AMP-equivalence question. The arXiv abstract explicitly states that most arguments in the paper were generated using GPT-6 Astra. The theorem and AI-role attribution are primary-source confirmed; the paper is new and has no independent external proof verification in the canonical record.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.2/10evidence 88/100
06 Sept 2026AI for ScienceBreakthrough

ChatGPT 5.6 Pro-obtained proofs improve hypergraph vertex-cover hardness bounds

Karthik C. S. and Dor Minzer present new multilayered PCP constructions yielding improved NP-hardness bounds for minimum vertex cover in uniform hypergraphs, including a tight k-epsilon hardness factor for k at least four without relying on the Unique Games Conjecture. The arXiv abstract states that the proofs were obtained using ChatGPT 5.6 Pro and subsequently rewritten by the authors. The results are primary-source confirmed but not independently reproduced or peer reviewed.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.1/10evidence 88/100
06 Sept 2026AI for ScienceBreakthrough

ChatGPT 5.6 Pro-suggested strategy helps settle critical Schouten-flow existence problem

Giovanni Catino and Carlo Mantegazza prove short-time existence and uniqueness for the critical Ricci-Bourguignon, or Schouten, flow, resolving the case left open by previous theory. The authors state that the strategy leading to the main argument was suggested during interactions with ChatGPT 5.6 Pro, after which they checked and developed the mathematics and take responsibility for the manuscript. The result is primary-source confirmed but not independently externally reproduced or peer reviewed.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.1/10evidence 89/100
04 Sept 2026AI for ScienceBreakthrough

Claude produces a complete machine-checked Lean formalization of Fermat's Last Theorem

Anthropic reports that a Claude Code-based multi-agent system produced the first complete computer-checked formalization of Fermat's Last Theorem in 11 days. The public artifact contains about 13 million lines of Lean and 29,511 theorem pages; its final theorem was checked by Lean, comparator and a second independently implemented Lean kernel. The result formalizes an existing theorem rather than discovering new mathematics, and the end-to-end run has not yet been independently reproduced by an external team.

PRIMARY CONFIRMEDANTHROPIC — FORMALIZING FERMAT'S LAST THEOREMPrimary source ↗
Relevance9.6/10evidence 96/100
03 Sept 2026AI for Science

Cherednik reports ChatGPT/Codex substantially contributed to proofs in instanton-superpolynomial work

In the v2 revision of 'Instanton slices and their superpolynomials', Ivan Cherednik explicitly attributes substantial mathematical and computational contributions to OpenAI ChatGPT/Codex, including a proof of topological invariance for instanton and motivic superpolynomials and a three-row mixed-characteristic comparison. The author says the work was cross-checked with independent implementations and his own calculations, but the 106-page preprint has not yet received external peer review or independent expert reproduction.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.9/10evidence 87/100
03 Sept 2026AI for Science

Google releases WeatherNext 3 with hourly high-resolution global forecasts

Google DeepMind and Google Research released WeatherNext 3 on September 3, 2026. The model ingests live geostationary satellite observations, refreshes forecasts hourly and produces selected surface variables at up to 5 km resolution. Google began integrating the forecasts across Search, the Gemini app, Google Maps, Google Maps Platform and Earth Engine. Google reports large precipitation-forecast improvements and points to Brightband's independent Operational WeatherBench; the provider benchmark magnitudes are not treated as independently reproduced in this assessment.

INDEPENDENTLY CONFIRMEDGOOGLE — INTRODUCING WEATHERNEXT 3Primary source ↗
Relevance8.7/10evidence 91/100
02 Sept 2026AI for ScienceBreakthrough

ChatGPT Sol 5.6 materially contributes to Cartan-convexity and butterfly-realization results

J. E. Pascoe develops Cartan convexity for self-adjoint free functions, proves extension results using universal direct sums and noncommutative Kraus-butterfly arguments, and establishes analogous results for graph embeddings. The arXiv comments state that the article was generated with ChatGPT Sol 5.6 and that the author reviewed the results and references. The mathematics and AI provenance are primary-source confirmed, while independent specialist reproduction is not established.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.0/10evidence 86/100
02 Sept 2026AI for Science

AI-assisted proof establishes the tight 4/3 correlation-gap bound for n=4

Arjun Ramachandra reports an AI-assisted proof that the pairwise-independent correlation gap for monotone submodular functions is universally bounded by 4/3 at n=4 and that the bound is tight. The paper says GPT 5.6 and Fable 5 helped explore the cone-certificate formulation and Bernstein representation; the author subsequently corrected and refined the strategy and verified 2,745 Bernstein coefficient systems computationally. The result remains an unreviewed preprint without independent external reproduction.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.8/10evidence 88/100
31 Aug 2026AI for ScienceBreakthrough

ChatGPT 5.6 helps find counterexample to the stable forking conjecture

James Freitag and Scott Mutchnik report a counterexample to the stable forking conjecture, a long-standing problem in model theory discussed since 1996. The authors state in the abstract that they found the counterexample using ChatGPT 5.6. The result is currently an arXiv preprint: the theorem and AI role are primary-source confirmed, but no peer review or independent mathematical reproduction was identified in this pass.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.1/10evidence 84/100
31 Aug 2026AI for Science

PaperGym trains small models to generate stronger research plans with rubric-centered feedback

PaperGym converts scientific papers into research-planning environments with separated questions and grading criteria, then combines rubric-conditioned self-distillation with rubric-reward training. Across Qwen3 models from 1.7B to 8B, the authors report five-benchmark average gains of 4.8 to 5.6 points and release the 20,000-instance corpus, benchmarks, models and training code.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.1/10evidence 82/100
31 Aug 2026AI for Science

AutoSciRub uses executable rubrics to improve autonomous research agents

AutoSciRub induces grounded, executable evaluation criteria before an autonomous research run and uses their results to target revisions. The authors report consistent ResearchClawBench gains across three model backbones and three agent harnesses, plus a 16.78-point average gain on a fixed 20-task AstaBench subset. The implementation is public.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.2/10evidence 80/100
31 Aug 2026AI for Science

Single-agent RL beats tree search on key chemistry tool-use metrics

Researchers report that one supervised-then-RL policy improves tool-selection and return metrics over CheMatAgent's hierarchical evolutionary tree search on ChemToolBench while using one model invocation per question. The result is material for scientific-agent efficiency, but the advantage is not uniform: on Llama 3.1 8B the search baseline retains a higher answer pass rate.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.1/10evidence 77/100
31 Aug 2026AI for Science

MedAgent-R1 sharply reduces fabricated citations in medical reasoning

MedAgent-R1 uses a faithfulness-gated reinforcement-learning reward to condition accuracy credit on evidence grounding. The authors report citation fabrication falling from 31.8% to 4.7% while evidence completeness rises to 82.6 and accuracy remains 75.1%. The result identifies a serious failure mode in outcome-only RL, but remains a preprint evaluation rather than clinical validation.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.3/10evidence 80/100
31 Aug 2026AI for Science

GPT-5.6 helps derive an explicit family of counterexamples to the Gaussian completely monotone conjecture

Jiayang Zou and Yihong Wu give a self-contained analytic construction of smooth, strictly log-concave counterexamples in every dimension. The authors report that they chose the ansatz and proof strategy, while GPT-5.6 Sol Pro identified the decisive parameter scaling and helped develop the exposition. The result extends earlier existence and discrete-counterexample work; it is not the first disproof of the conjecture.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.6/10evidence 84/100
31 Aug 2026AI for Science

AI-assisted heterogeneous recursion improves the best known lower bound for the Shannon capacity of C7

Ravi Tandon reports a heterogeneous refinement of recursive zero-error Shannon-capacity constructions, producing an explicit independent set in the 500th strong power of the seven-cycle and a lower bound Theta(C7) >= 3.25883262..., improving the previous best bound. The paper presents the work as AI-assisted. No independent external reproduction of the final construction was verified.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.7/10evidence 88/100
31 Aug 2026AI for Science

Cubic-root Gaussian approximation falsifies an n^-1/4 rate conjecture

A new proof establishes an n^-1/3 high-dimensional Gaussian-approximation rate under unrestricted covariance in the polynomial-dimensional regime, falsifying a 2023 n^-1/4 near-optimality conjecture. The authors state that ChatGPT 5.6 Pro generated the initial proof attempt, which they corrected and rewrote, and provide a Lean formalization. The theorem is substantial, but the AI role was assistive rather than autonomous and independent replication is absent.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.6/10evidence 83/100
30 Aug 2026AI for Science

GPT Pro produces 18 verified counterexamples across 14 mathematical cases

A 35-page preprint and open repository consolidate 18 refuted statements across 14 cases found largely with GPT Pro. Sixteen recompute with exact rational arithmetic and two with rigorous interval enclosures; the author supplies human-readable proofs and machine-verifiable artifacts.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.9/10evidence 86/100
28 Aug 2026AI for Science

ChatGPT 5.6 Sol Pro helped break the Bethe barrier for deterministic permanent approximation

A Stanford preprint gives a deterministic polynomial-time approximation for the permanent of every nonnegative matrix with factor c^n for an absolute c below sqrt(2), improving the previous universal Bethe-based base. Nima Anari supplied and guided the high-level plan; the paper states that ChatGPT 5.6 Sol Pro proposed three central proof strategies. An accompanying Lean 4 development formalizes the proof and complete algorithm.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.4/10evidence 74/100
27 Aug 2026AI for Science

Anthropic launches Model Hardware Standard; Claude develops and tunes a quantum-laser controller

Anthropic opened a research preview of MHS, a model-agnostic interface for agents to operate programmable physical equipment. In a QuEra pilot, Claude used MHS to develop a deterministic laser-relock controller validated on 700 induced disturbances and to tune a live laser system in a bounded, unattended loop.

PRIMARY CONFIRMEDANTHROPIC — PREVIEWING THE MODEL HARDWARE STANDARDPrimary source ↗
Relevance8.9/10evidence 84/100
27 Aug 2026AI for Science

AI co-developed proofs of Fröberg’s conjecture for quintic and septic slices in four variables

A revised preprint proves Fröberg’s predicted Hilbert series for every generator count in the equal-degree four-variable cases d=5 and d=7. The authors disclose a generative-AI workflow using GPT-5.6 Sol, Claude Fable 5 and Grok 4.6.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.5/10evidence 74/100
27 Aug 2026AI for Science

Claude Fable 5 helps find counterexample for entanglement-of-formation support functionals

A Steklov Mathematical Institute preprint gives an explicit two-qubit counterexample showing that a global supporting affine functional for entanglement of formation need not exist at a degenerate state; the authors report Claude Fable 5 quickly found the crucial state.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.3/10evidence 76/100
27 Aug 2026AI for Science

GPT-5.6 Sol designs near-optimal algorithms across three operations-research domains

An NYU Stern preprint reports that a single untuned GPT-5.6 Sol query produced reusable algorithms that matched or outperformed the best reported methods on almost all evaluated inventory, queueing and assortment instances, including frozen holdout evaluations.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.7/10evidence 81/100
27 Aug 2026AI for Science

GPT-5.6 Sol supplies decisive positivity certificate in dual futile-cycle proof

A Leipzig University preprint proves that a sequential distributive dual futile-cycle model can admit Hopf bifurcations under parameter-rich kinetics but not under mass-action kinetics; the author reports that GPT-5.6 Sol produced the decisive positivity certificate in one query.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.6/10evidence 78/100
27 Aug 2026AI for Science

ChatGPT 5.6 Sol-assisted paper answers higher-order truth question negatively

Yiqi Xu and Lingyuan Ye show that the free Heyting algebra on two generators, and hence on any finite number greater than two, cannot occur as the lattice of subterminal objects of an elementary topos, answering the stated question negatively. The abstract explicitly says the mathematical results were obtained with the help of ChatGPT 5.6 Sol. The paper was first submitted August 27 and revised September 3; the recovered v2 signal is primary-source confirmed but not independently reproduced.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.9/10evidence 87/100