OBSERVATORY DEGRADED

Semantic review is behind discovery. Evidence and Pulse may be incomplete until the backlog is cleared.

2 material first-party/frontier/open-problem candidate(s) have waited more than 6 hours for semantic review.

0 critical overdue · 2 material overdue
Last semantic import 19 Sept, 04:12

EVIDENCE REGISTER

The evidence archive

A chronological register of published signals. Source class, verification state, evidence quality and relevance remain visible before interpretation.

PUBLISHED RECORDS

222 matching evidence records

1–25 / 222
18 Sept 2026Intelligence

Anthropic and Accenture launch embedded frontier-AI safety evaluation partnership

Anthropic and Accenture announced a partnership to place dedicated external evaluators alongside Anthropic teams for model evaluation, red-teaming, alignment assessment and safeguard testing. Reporting on the announcement says each company expects to invest at least $1 billion over five years. The arrangement is significant as evaluation infrastructure and governance, but does not itself establish improved model capability, safety performance, or independent reproduction of any model claim.

INDEPENDENTLY CONFIRMEDANTHROPIC — PARTNERING WITH ACCENTURE ON EMBEDDED EVALUATIONPrimary source ↗
Relevance8.4/10evidence 92/100
17 Sept 2026AI for Science

GPT-6 Astra-assisted proof settles strong secretary conjecture for linear matroids

Bérczi, Dughmi, Livanos, Soto and Verdugo prove a 1/e guarantee for the matroid secretary problem on linear matroids and state that the main proof was obtained in a conversation with ChatGPT-6 Astra. A separately authored concurrent preprint by Abdi, Banihashem, Hajiaghayi and Mittal independently proves the same linear-matroid result via an essentially identical approach.

INDEPENDENTLY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.8/10evidence 96/100
17 Sept 2026AI for Science

GPT-6 Astra-assisted proof establishes expectation form of BHM conjecture

Yinfeng Zhu proves that paths maximize the expected range of integer-valued graph homomorphisms among connected bipartite graphs of fixed order, establishing the expectation form of the Benjamini-Häggström-Mossel conjecture and deriving the Loebl-Nešetřil-Reed conjecture as a corollary. The author states that the proof was obtained through interaction with GPT-6 Astra and that the main results were formalized and checked in Lean 4.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.8/10evidence 94/100
17 Sept 2026AI for Science

GPT-5.6 Sol Ultra-assisted work sharpens trickle-down spectral-gap theorem

Xiaoyu Chen and Kuikui Liu give streamlined Bochner-method proofs of trickle-down spectral-gap results and quantitatively strengthen the Leake-Oveis Gharan theorem, resolving an open question. They state that the proofs were developed through interaction with GPT-5.6 Sol Ultra and note that Guo and Zhang independently obtained the same strengthening with a very similar argument, also found using GPT-5.6 Sol Ultra.

INDEPENDENTLY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.7/10evidence 95/100
17 Sept 2026IntelligenceBreakthrough

Anthropic reports Claude leads 26% of its measured AI R&D work

Anthropic's R&D Automation Index reports that as of August 2026 Claude leads 26% of measured AI R&D work from high-level prompts under human supervision, while more than 90% is at least human-AI collaborative. Anthropic explicitly reports no measured subset at full autonomy and notes methodological limitations including use of its own models as judges.

PRIMARY CONFIRMEDANTHROPIC INSTITUTE — MEASUREMENTS FOR UNDERSTANDING THE PACE OF AI DEVELOPMENT INSIDE FRONTIER LABSPrimary source ↗
Relevance9.7/10evidence 86/100
17 Sept 2026AI for Science

Claude optimizes more than 30 open-source biomolecular models

Anthropic reports that Claude, supervised by two technical staff, optimized more than 30 open-source models across structure prediction, protein design, protein language modeling and genomics in under four weeks, averaging roughly 4x speed-ups with minimal precision loss and creating a low-memory mode for much larger biomolecular systems. Optimized code was released publicly.

PRIMARY CONFIRMEDANTHROPIC — HOW CLAUDE IS UPLIFTING BIOMOLECULAR MODELINGPrimary source ↗
Relevance8.9/10evidence 86/100
17 Sept 2026Intelligence

Anthropic launches Life Sciences Verification Program beta

Anthropic opened applications for a beta program giving verified life-science teams access to Mythos, Opus and Sonnet models under more permissive biology safeguards, with separate Standard Use and project-specific High-risk Use grants and continuous scope monitoring.

ANNOUNCEDANTHROPIC — INTRODUCING THE LIFE SCIENCES VERIFICATION PROGRAMPrimary source ↗
Relevance8.0/10evidence 88/100
16 Sept 2026AI for Science

ChatGPT 5.6 Sol finds proofs for stronger Gorenstein-polytope decomposition results

Johannes Knupfer and Benjamin Nill prove a free-join decomposition theorem for Gorenstein polytopes that significantly strengthens prior work and prove the expected-degree characterization of the stringy E-polynomial, resolving a conjecture of Batyrev and Nill. The authors state that proofs of the general results were found via ChatGPT 5.6 Sol, while an extremal case had a prior human proof.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.6/10evidence 91/100
16 Sept 2026AI for Science

GPT-5.6 Sol-assisted work establishes degree-free spectral independence for log-concave Holant measures

Xiaoyu Chen, Zejia Chen and Xinyuan Zhang establish a degree-independent spectral-independence bound for log-concave Holant problems and derive improved Glauber-dynamics relaxation bounds. The authors explicitly state that the main proof ideas were found using GPT-5.6 Sol.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.6/10evidence 92/100
16 Sept 2026AI for Science

ChatGPT-6 Astra assists counterexample to proposed relay-channel capacity characterization

Chun Hei Michael Shiu constructs a binary relay-channel counterexample whose achievable scheme exceeds a recently proposed capacity characterization, disproving that characterization for general relay channels. The paper explicitly states that the counterexample was obtained with assistance from ChatGPT-6 Astra and that the proofs were simplified through human effort.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.1/10evidence 89/100
16 Sept 2026AI for Science

ChatGPT 5.6-assisted work improves register-program bounds for catalytic computing

Antoine Vinciguerra proves the optimality of four input accesses for passive-output register programs of degree greater than three and constructs improved four-access register programs using fewer registers, yielding improved catalytic-streaming and matrix-powering trade-offs. The paper states that generalizing the lower-bound methods and constructing the uniform family were developed with assistance from ChatGPT 5.6.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.5/10evidence 92/100
15 Sept 2026AI for ScienceBreakthrough

Independent mathematicians present GPT-6 Astra proof of the Erdős–Sós conjecture

Oliver Riordan and Alex Scott state that the Erdős–Sós conjecture was recently proved by GPT-6 Astra using a surprising argument and publish a simplified version, also determining extremal graphs and proving a related conjecture. David R. Wood separately publishes an exposition of the proof, likewise attributing its discovery to GPT-6 Astra. The multiple expert rewrites provide unusually strong independent validation of the AI-origin claim.

INDEPENDENTLY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.9/10evidence 97/100
15 Sept 2026AI for Science

ChatGPT 5.6 Sol-assisted work resolves cluster-deletion complexity question

Nicola Galesi, Tony Huynh, Arnaud Patey and Fariba Ranjbar prove Cluster Deletion is NP-complete on permutation graphs, answering an open question of Konstantinidis and Papadopoulos. The authors state that Patey found the proof with help from ChatGPT 5.6 Sol; the model-generated proof contained errors and was completely rewritten by the authors.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance8.5/10evidence 87/100
15 Sept 2026IntelligenceBreakthrough

Google releases Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking for near-real-time multimodal voice agents. The models support background tool/API calls while conversation continues, visual grounding and, for Extended Thinking, simultaneous reasoning and speech. Artificial Analysis independently lists Gemini 3.8 Live Extended Thinking at 82.6 on its Speech-to-Speech Index and 68.6% on τ-Voice, while Gemini 3.8 Live scores 76.0 on the same aggregate index.

INDEPENDENTLY CONFIRMEDGOOGLE — INTRODUCING GEMINI 3.8 LIVE AND 3.8 LIVE EXTENDED THINKINGPrimary source ↗
Relevance9.3/10evidence 96/100
14 Sept 2026AI for ScienceBreakthrough

ChatGPT 6 Astra supplies key strategies resolving 2D Le simplex conjecture

Chong Gu proves that triangles minimize the Monge–Ampère eigenvalue among bounded planar convex domains of fixed area, resolving the two-dimensional case of Le's simplex conjecture. The author states that the main results were obtained through a series of chats with ChatGPT 6 Astra and that the key strategies came from the model, after which he reworked and rewrote the article.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.3/10evidence 90/100
13 Sept 2026AI for ScienceBreakthrough

GPT-5.6 Pro-assisted proof closes Ryser (4,2) case

Patrick White proves that every 4-partite 4-uniform hypergraph with matching number two has vertex-cover number at most six, confirming Tuza's unpublished 1979 claim and closing the (r,ν)=(4,2) case of Ryser's conjecture. The author reports the proof was found through four rounds of GPT-5.6 Pro and that structural claims were checked by exact MILP and brute-force computation.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.2/10evidence 94/100
12 Sept 2026AI for ScienceBreakthrough

GPT-6 Astra-assisted proof settles Conway subprime-closure growth conjecture

Romain Popescu proves that the growth ratio of Conway's subprime closure converges to the golden ratio, settling the conjecture of Caragiu, Vicol and Zaki. The paper states that the underlying proof was constructed with algorithmic assistance from GPT-6 Astra and that its correctness was formally verified in Lean 4.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.0/10evidence 94/100
12 Sept 2026AI for Science

GPT-6 Astra assists simplified proof of Rockafellar sum-conjecture failure

Radu Ioan Bot presents a simplified counterexample to Rockafellar's sum conjecture on a concrete Banach-space setting, building on a construction of Weifeng Yang. The paper says the example was developed with GPT-6 Astra assistance and explicitly cautions that it is a simplification and explanation, not an independent counterexample mechanism.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance7.8/10evidence 86/100
11 Sept 2026AI for ScienceBreakthrough

ChatGPT Sol 5.6 supplies core proof work settling undecidability of Medvedev logic

Rodrigo Nicolau Almeida and Søren Brinck Knudstorp prove Medvedev's logic of finite problems is undecidable, settling a longstanding open problem, and obtain related results for Skvortsov's logic. The paper states that the core idea and technical work of the undecidability proof were obtained using ChatGPT Sol 5.6 and formally verified in Lean by Claude Opus 5.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.5/10evidence 95/100
10 Sept 2026AI for ScienceBreakthrough

GPT-6 Astra-obtained proof settles core existence in approval-based committee elections

Patrick Becker, Matthias Greger and Dominik Peters prove that every approval-based multiwinner election admits a core-stable committee and that a Hare-core committee can be found in polynomial time, settling a central open question in the field. The arXiv record explicitly states that the proof was obtained with GPT-6 Astra. The paper also points to a Lean formalization of the core-existence theorem. This provides strong artifact-level support, but the result remains a new author-controlled preprint without independent external expert reproduction.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.5/10evidence 94/100
10 Sept 2026AI for ScienceBreakthrough

GPT-6 Astra obtains counterexample disproving the wrapping number conjecture

Qiuyu Ren exhibits an annular knot with wrapping number four whose Kauffman bracket has annular degree at most two, disproving the wrapping number conjecture. The arXiv comments explicitly state that the main result was obtained by GPT-6 Astra. The theorem and provenance are primary-source confirmed, but the four-page preprint has not yet received independent expert verification or reproduction.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.3/10evidence 87/100
10 Sept 2026AI for ScienceBreakthrough

ChatGPT Pro 6.0-assisted paper settles a conjecture on II1 factors of Fuchsian groups

Dimitri Shlyakhtenko proves that the von Neumann algebra of the fundamental group of any closed orientable surface of genus at least two is a free group factor, and extends the conclusion to arbitrary finitely generated torsion-free non-elementary discrete subgroups of PSL2(R), settling a conjecture of de la Harpe and Voiculescu. The arXiv abstract explicitly states that the result was obtained using OpenAI's ChatGPT Pro 6.0. The mathematical claim and AI-role attribution are primary-source confirmed, but the proof is a new preprint without independent expert reproduction or peer review.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.3/10evidence 88/100
10 Sept 2026IntelligenceBreakthrough

Anthropic reports real-world Claude misuse across cyber, surveillance, weapons and biological cases

Anthropic's September 2026 Threat Intelligence report describes operations disrupted between December 2025 and August 2026 involving Claude in cyber operations, surveillance, influence operations, scams, conventional-weapons work, biological misuse and illicit distillation. Reuters and AP independently reported major categories and examples from the disclosure. The case studies are real-world misuse evidence, but most underlying attribution and technical detail comes from Anthropic and is not independently reproduced.

INDEPENDENTLY CONFIRMEDANTHROPIC — DETECTING AND COUNTERING MISUSE OF AI: SEPTEMBER 2026Primary source ↗
Relevance9.3/10evidence 94/100
10 Sept 2026IntelligenceBreakthrough

Anthropic reports frontier AI reaching scarce-expert performance on some targeting and weapons tasks

Anthropic's Frontier Red Team released evaluations of tactical intelligence targeting and conventional-weapons development. Anthropic reports that on some tasks frontier models could perform work historically limited to scarce, highly trained human experts, including geolocating people from fragmentary information and engineering tasks related to drones and moving targets. The evaluation is provider-run and is not independently reproduced, so capability magnitudes remain first-party claims.

PRIMARY CONFIRMEDANTHROPIC — MEASURING TACTICAL INTELLIGENCE TARGETING AND CONVENTIONAL WEAPONS CAPABILITIES OF AI MODELSPrimary source ↗
Relevance9.1/10evidence 91/100
09 Sept 2026AI for ScienceBreakthrough

LLM-assisted example resolves zero-private-capacity superactivation problem

Chengkai Zhu and Xin Wang resolve a longstanding quantum-information question by constructing two channels that each have zero private capacity but jointly achieve positive private communication. The paper states that the initial activation example was identified through interactions with large language models and that the result has been formalized in Lean 4. The theorem has strong author-provided formal-verification support but has not yet been independently externally reproduced.

PRIMARY CONFIRMEDPREPRINTPrimary source ↗
Relevance9.5/10evidence 94/100
Page 1 / 9
← PreviousNext →