OBSERVATORY DEGRADED

Semantic review is behind discovery. Evidence and Pulse may be incomplete until the backlog is cleared.

2 material first-party/frontier/open-problem candidate(s) have waited more than 6 hours for semantic review.

0 critical overdue · 2 material overdue
Last semantic import 19 Sept, 04:12
← Back to Pulse
DAILY REVIEW23 Aug 2026

Daily Pulse — 23 Aug 2026

No frontier-model capability jump in the last 24h. Material signals are a first-party Grok multi-agent engineering loop reaching physical manufacture, reported >15% AI-server price increases driven by memory costs, and new humanoid locomotion records. The open-problem audit recovered six missed qualifying radar records; AI for Science maturity and ASI forecast remain unchanged.

ANALYSIS

Methodology note: Historical index values on this Pulse are normalized to the current methodology for comparability. The original as-published record remains unchanged in the canonical archive.

Verdict: no new frontier-model capability ceiling mover in the last 24 hours, so Intelligence, AI for Science, Robotics, ASI Readiness and the ASI forecast remain unchanged. Three fresh signals are material enough to record: (1) a first-party Grok multi-agent engineering demonstration that converted four reference images into a physics simulator, ran 47 fracture experiments, revised its hypothesis, optimized designs and produced 3D-printed physical objects; (2) a reported >15% increase in prices for many NVIDIA AI-server configurations driven by memory costs, affecting Vera Rubin and Grace Blackwell systems shipping in early 2027; and (3) independently reported humanoid sprint and standing-jump records at the Beijing World Humanoid Robot Games.

Frontier intelligence / release coverage: two-pass release checks across OpenAI, Anthropic, Google DeepMind, xAI, DeepSeek and Qwen/Alibaba found no verified new flagship/frontier release after DeepSeek V4-Flash-Vision-Exp. OpenAI first-party release surfaces remain on GPT-5.6; Google DeepMind remains on Gemini 3.7 Flash as its latest frontier-line update. Absence from any single indexed page was not treated as proof.

Agents / autonomy / AI-R&D: Markus Buehler reports a Grok team-of-agents workflow that autonomously inferred structural principles from four images, built an executable physics simulator, ran 47 fracture experiments, rejected an initial hypothesis, selected designs, sliced them and manufactured them by 3D printing. Evidence is first-party and the demonstration has not been independently replicated. Relevance 7.8/10, evidence 78/100, impact_asi 3/10. It is a meaningful engineering-loop signal but not enough to move the autonomy or recursive-R&D anchor.

AI for Science / Open Problem Radar: six qualifying records missed by prior screening were recovered and marked discovered_in_audit. The strongest AI-material additions are the Gaussian boson sampling hiding conjecture (full claim, assisted, significance 7.7, evidence 88) and the Hilbert-space inverse generator problem (full negative solution, assisted, significance 7.1, evidence 84). The audit also recovered the Bray volume conjecture (human-only, 6.5), the selfinjective Snashall-Solberg case (human-only, 5.8), Strong Graph Reconstruction / BBF deck-overlap bound (co-developed, 5.8), and the ordinary-vs-strong Kreiss separation question (assisted, 5.4). The rank-30 elliptic-curve record already present in the previous Pulse received a provenance update to assisted after primary leaderboard attribution to Claude with Levent Alpoge and Ava Howell; the division of labor remains insufficiently documented to move maturity. No independently verified new closure was found today.

Discovery–Verification Gap: the conservative 90-day qualifying discovery count rises from 3 to 5 after adding the GBS hiding and inverse-generator results; verified remains 2. Discovery rate 1.69/month, verification rate 0.68/month, DVG 2.5x. Coverage remains partial.

Robotics: Tiangong Ultra completed the official 100 m final in 9.39 s and the event also produced a 2.88 m standing-jump record; independent reporting confirms the results. These are strong dynamic-locomotion signals, not evidence of general-purpose autonomy, task generalization or low intervention rates. Relevance 7.2/10, evidence 92/100, impact_asi 0.2/10. Robotics stays 47.4.

Compute & Infrastructure: Reuters reports, citing Bloomberg, that some major NVIDIA customers have been told AI-server prices will rise by more than 15% in many cases because of memory costs, including Vera Rubin and Grace Blackwell systems expected to ship in early 2027. Reuters could not independently verify the report and NVIDIA had not commented. Relevance 7.0/10, evidence 68/100, impact_asi 0. This strengthens memory/HBM economics as a bottleneck but does not change Compute maturity.

Synthetic Biology, Energy/Fusion, Advanced Materials, Longevity and BCI/HMI: no new 24h evidence crossed a maturity anchor. SynBio remains 66.8, Energy/Fusion 46.8, Materials 62.4, Longevity 36.3 and BCI/HMI 42.8.

Benchmark Observatory: METR Time Horizon remains on TH 1.1 with the public page last updated 8 May 2026 and explicit warning that measurements above 16 hours are unreliable with the current task suite. Epoch AI shows FrontierMath data updated 21 Aug; no new verified point requiring a readiness change was found. Benchmark saturation/exposure and scaffold comparability remain measurement bottlenecks.

Readiness: Intelligence 63.7; AI for Science 59.0; Robotics 47.4. ASI Readiness 59.03 (delta 0.00). Singularity Readiness 53.19 (delta 0.00). Rolling ASI forecast unchanged: 1y 2%, 2y 6%, 3y 12%, 5y 30%, 10y 81%; central estimate 2032–33.

Bottom line: capability did not jump today, but the engineering loop is worth watching because it joins a growing class of systems that can carry an objective across reasoning, simulation, iteration and a physical endpoint. The stronger constraint remains reliability and independent replication. On infrastructure, memory cost pressure is a reminder that scaling intelligence is increasingly constrained by the whole compute stack rather than accelerators alone.

Audit: six Open Problem Radar misses were recovered and corrected; no additional >=7/10 24–48h event absent from the database was found after the second pass. Audit remains partial because exhaustive enumeration of the full arXiv daily feed cannot be guaranteed on a Sunday/weekend window.