07 Sept 2026AI for ScienceBreakthrough
GPT-6 Astra generates most arguments in result on AMP and low-degree polynomial equivalence
Zhangsong Li proves an almost-sharp equivalence between approximate message passing and growing-degree polynomial estimation for the Bernoulli rank-one planted-submatrix setting, resolving the Bernoulli rank-one case of a growing-degree AMP-equivalence question. The arXiv abstract explicitly states that most arguments in the paper were generated using GPT-6 Astra. The theorem and AI-role attribution are primary-source confirmed; the paper is new and has no independent external proof verification in the canonical record.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance9.2/10evidence 88/100
06 Sept 2026AI for ScienceBreakthrough
ChatGPT 5.6 Pro-obtained proofs improve hypergraph vertex-cover hardness bounds
Karthik C. S. and Dor Minzer present new multilayered PCP constructions yielding improved NP-hardness bounds for minimum vertex cover in uniform hypergraphs, including a tight k-epsilon hardness factor for k at least four without relying on the Unique Games Conjecture. The arXiv abstract states that the proofs were obtained using ChatGPT 5.6 Pro and subsequently rewritten by the authors. The results are primary-source confirmed but not independently reproduced or peer reviewed.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance9.1/10evidence 88/100
06 Sept 2026AI for ScienceBreakthrough
ChatGPT 5.6 Pro-suggested strategy helps settle critical Schouten-flow existence problem
Giovanni Catino and Carlo Mantegazza prove short-time existence and uniqueness for the critical Ricci-Bourguignon, or Schouten, flow, resolving the case left open by previous theory. The authors state that the strategy leading to the main argument was suggested during interactions with ChatGPT 5.6 Pro, after which they checked and developed the mathematics and take responsibility for the manuscript. The result is primary-source confirmed but not independently externally reproduced or peer reviewed.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance9.1/10evidence 89/100
04 Sept 2026AI for ScienceBreakthrough
Claude produces a complete machine-checked Lean formalization of Fermat's Last Theorem
Anthropic reports that a Claude Code-based multi-agent system produced the first complete computer-checked formalization of Fermat's Last Theorem in 11 days. The public artifact contains about 13 million lines of Lean and 29,511 theorem pages; its final theorem was checked by Lean, comparator and a second independently implemented Lean kernel. The result formalizes an existing theorem rather than discovering new mathematics, and the end-to-end run has not yet been independently reproduced by an external team.
PRIMARY CONFIRMEDANTHROPIC — FORMALIZING FERMAT'S LAST THEOREMPrimary source ↗ Relevance9.6/10evidence 96/100
03 Sept 2026AI for Science
Cherednik reports ChatGPT/Codex substantially contributed to proofs in instanton-superpolynomial work
In the v2 revision of 'Instanton slices and their superpolynomials', Ivan Cherednik explicitly attributes substantial mathematical and computational contributions to OpenAI ChatGPT/Codex, including a proof of topological invariance for instanton and motivic superpolynomials and a three-row mixed-characteristic comparison. The author says the work was cross-checked with independent implementations and his own calculations, but the 106-page preprint has not yet received external peer review or independent expert reproduction.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.9/10evidence 87/100
03 Sept 2026AI for Science
Google releases WeatherNext 3 with hourly high-resolution global forecasts
Google DeepMind and Google Research released WeatherNext 3 on September 3, 2026. The model ingests live geostationary satellite observations, refreshes forecasts hourly and produces selected surface variables at up to 5 km resolution. Google began integrating the forecasts across Search, the Gemini app, Google Maps, Google Maps Platform and Earth Engine. Google reports large precipitation-forecast improvements and points to Brightband's independent Operational WeatherBench; the provider benchmark magnitudes are not treated as independently reproduced in this assessment.
INDEPENDENTLY CONFIRMEDGOOGLE — INTRODUCING WEATHERNEXT 3Primary source ↗ Relevance8.7/10evidence 91/100
02 Sept 2026AI for ScienceBreakthrough
ChatGPT Sol 5.6 materially contributes to Cartan-convexity and butterfly-realization results
J. E. Pascoe develops Cartan convexity for self-adjoint free functions, proves extension results using universal direct sums and noncommutative Kraus-butterfly arguments, and establishes analogous results for graph embeddings. The arXiv comments state that the article was generated with ChatGPT Sol 5.6 and that the author reviewed the results and references. The mathematics and AI provenance are primary-source confirmed, while independent specialist reproduction is not established.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance9.0/10evidence 86/100
02 Sept 2026AI for Science
AI-assisted proof establishes the tight 4/3 correlation-gap bound for n=4
Arjun Ramachandra reports an AI-assisted proof that the pairwise-independent correlation gap for monotone submodular functions is universally bounded by 4/3 at n=4 and that the bound is tight. The paper says GPT 5.6 and Fable 5 helped explore the cone-certificate formulation and Bernstein representation; the author subsequently corrected and refined the strategy and verified 2,745 Bernstein coefficient systems computationally. The result remains an unreviewed preprint without independent external reproduction.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.8/10evidence 88/100
31 Aug 2026AI for ScienceBreakthrough
ChatGPT 5.6 helps find counterexample to the stable forking conjecture
James Freitag and Scott Mutchnik report a counterexample to the stable forking conjecture, a long-standing problem in model theory discussed since 1996. The authors state in the abstract that they found the counterexample using ChatGPT 5.6. The result is currently an arXiv preprint: the theorem and AI role are primary-source confirmed, but no peer review or independent mathematical reproduction was identified in this pass.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance9.1/10evidence 84/100
31 Aug 2026AI for Science
PaperGym trains small models to generate stronger research plans with rubric-centered feedback
PaperGym converts scientific papers into research-planning environments with separated questions and grading criteria, then combines rubric-conditioned self-distillation with rubric-reward training. Across Qwen3 models from 1.7B to 8B, the authors report five-benchmark average gains of 4.8 to 5.6 points and release the 20,000-instance corpus, benchmarks, models and training code.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.1/10evidence 82/100
31 Aug 2026AI for Science
AutoSciRub uses executable rubrics to improve autonomous research agents
AutoSciRub induces grounded, executable evaluation criteria before an autonomous research run and uses their results to target revisions. The authors report consistent ResearchClawBench gains across three model backbones and three agent harnesses, plus a 16.78-point average gain on a fixed 20-task AstaBench subset. The implementation is public.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.2/10evidence 80/100
31 Aug 2026AI for Science
Single-agent RL beats tree search on key chemistry tool-use metrics
Researchers report that one supervised-then-RL policy improves tool-selection and return metrics over CheMatAgent's hierarchical evolutionary tree search on ChemToolBench while using one model invocation per question. The result is material for scientific-agent efficiency, but the advantage is not uniform: on Llama 3.1 8B the search baseline retains a higher answer pass rate.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.1/10evidence 77/100
31 Aug 2026AI for Science
MedAgent-R1 sharply reduces fabricated citations in medical reasoning
MedAgent-R1 uses a faithfulness-gated reinforcement-learning reward to condition accuracy credit on evidence grounding. The authors report citation fabrication falling from 31.8% to 4.7% while evidence completeness rises to 82.6 and accuracy remains 75.1%. The result identifies a serious failure mode in outcome-only RL, but remains a preprint evaluation rather than clinical validation.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.3/10evidence 80/100
31 Aug 2026AI for Science
GPT-5.6 helps derive an explicit family of counterexamples to the Gaussian completely monotone conjecture
Jiayang Zou and Yihong Wu give a self-contained analytic construction of smooth, strictly log-concave counterexamples in every dimension. The authors report that they chose the ansatz and proof strategy, while GPT-5.6 Sol Pro identified the decisive parameter scaling and helped develop the exposition. The result extends earlier existence and discrete-counterexample work; it is not the first disproof of the conjecture.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.6/10evidence 84/100
31 Aug 2026AI for Science
AI-assisted heterogeneous recursion improves the best known lower bound for the Shannon capacity of C7
Ravi Tandon reports a heterogeneous refinement of recursive zero-error Shannon-capacity constructions, producing an explicit independent set in the 500th strong power of the seven-cycle and a lower bound Theta(C7) >= 3.25883262..., improving the previous best bound. The paper presents the work as AI-assisted. No independent external reproduction of the final construction was verified.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.7/10evidence 88/100
31 Aug 2026AI for Science
Cubic-root Gaussian approximation falsifies an n^-1/4 rate conjecture
A new proof establishes an n^-1/3 high-dimensional Gaussian-approximation rate under unrestricted covariance in the polynomial-dimensional regime, falsifying a 2023 n^-1/4 near-optimality conjecture. The authors state that ChatGPT 5.6 Pro generated the initial proof attempt, which they corrected and rewrote, and provide a Lean formalization. The theorem is substantial, but the AI role was assistive rather than autonomous and independent replication is absent.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.6/10evidence 83/100
30 Aug 2026AI for Science
GPT Pro produces 18 verified counterexamples across 14 mathematical cases
A 35-page preprint and open repository consolidate 18 refuted statements across 14 cases found largely with GPT Pro. Sixteen recompute with exact rational arithmetic and two with rigorous interval enclosures; the author supplies human-readable proofs and machine-verifiable artifacts.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.9/10evidence 86/100
28 Aug 2026AI for Science
ChatGPT 5.6 Sol Pro helped break the Bethe barrier for deterministic permanent approximation
A Stanford preprint gives a deterministic polynomial-time approximation for the permanent of every nonnegative matrix with factor c^n for an absolute c below sqrt(2), improving the previous universal Bethe-based base. Nima Anari supplied and guided the high-level plan; the paper states that ChatGPT 5.6 Sol Pro proposed three central proof strategies. An accompanying Lean 4 development formalizes the proof and complete algorithm.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.4/10evidence 74/100
27 Aug 2026AI for Science
Anthropic launches Model Hardware Standard; Claude develops and tunes a quantum-laser controller
Anthropic opened a research preview of MHS, a model-agnostic interface for agents to operate programmable physical equipment. In a QuEra pilot, Claude used MHS to develop a deterministic laser-relock controller validated on 700 induced disturbances and to tune a live laser system in a bounded, unattended loop.
PRIMARY CONFIRMEDANTHROPIC — PREVIEWING THE MODEL HARDWARE STANDARDPrimary source ↗ Relevance8.9/10evidence 84/100
27 Aug 2026AI for Science
GPT-5.6 Sol designs near-optimal algorithms across three operations-research domains
An NYU Stern preprint reports that a single untuned GPT-5.6 Sol query produced reusable algorithms that matched or outperformed the best reported methods on almost all evaluated inventory, queueing and assortment instances, including frozen holdout evaluations.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.7/10evidence 81/100
27 Aug 2026AI for Science
ChatGPT 5.6 Sol-assisted paper answers higher-order truth question negatively
Yiqi Xu and Lingyuan Ye show that the free Heyting algebra on two generators, and hence on any finite number greater than two, cannot occur as the lattice of subterminal objects of an elementary topos, answering the stated question negatively. The abstract explicitly says the mathematical results were obtained with the help of ChatGPT 5.6 Sol. The paper was first submitted August 27 and revised September 3; the recovered v2 signal is primary-source confirmed but not independently reproduced.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.9/10evidence 87/100
27 Aug 2026AI for Science
AgentFold runs closed-loop agentic search over protein-folding model designs
A multi-university preprint reports a multi-agent system that proposed, implemented, debugged, trained and evaluated roughly 80 ESMFold-derived code variants, outperforming matched Codex-proposal and random-search controls on a development benchmark.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance8.4/10evidence 79/100
27 Aug 2026AI for Science
Google extends Co-Scientist into closed-loop real-world scientific research
A Google-led preprint extends Gemini-based Co-Scientist from hypothesis generation into execution-grounded workflows spanning experiment planning, laboratory interaction, real-data validation and manuscript generation across materials science, biology and computer science.
PRIMARY CONFIRMEDPREPRINTPrimary source ↗ Relevance9.1/10evidence 80/100
27 Aug 2026AI for Science
Gemini 3.7 Flash reproduces three open-problem results with Antigravity Teamwork
Google Antigravity reports seven notable mathematics and theoretical-computer-science results from Teamwork. The detailed first-party report states the seven results were obtained with Gemini 3.1 Pro and that Problems 1, 3 and 4 were independently rerun within the Teamwork system using Gemini 3.7 Flash. The same framework operates autonomously over hours or days, retains failed approaches and verifier findings across rounds, and also built a RISC-V simulator validated against external timing oracles.
PRIMARY CONFIRMEDGOOGLE ANTIGRAVITYPrimary source ↗ Relevance8.8/10evidence 88/100
26 Aug 2026AI for Science
UCL and UCLH report first live AI decision support during brain-tumour surgery
A UCL-developed system analysed endoscopic video during pituitary-tumour surgery and highlighted critical anatomy in real time while the neurosurgeon retained full control; the tumour was removed and the patient’s vision improved.
INDEPENDENTLY CONFIRMEDUCLPrimary source ↗ Relevance8.3/10evidence 86/100