1. OpenAI's Astra Solves Ten Decade-Old Math Problems
On August 1, OpenAI published machine-checked Lean 4 proofs generated by Astra, its unreleased next-generation model, resolving ten open problems in mathematics and theoretical computer science that had stood for at least a decade. The headline result is the first explicit construction of a non-sofic group, closing a question in group theory posed by Mikhail Gromov in 1999. OpenAI estimates the compute cost for all ten solutions at roughly $2,000 and released the proof files and reasoning transcripts publicly on GitHub under Apache 2.0.
2. EU AI Act Transparency Rules Take Effect Across the Bloc
Starting August 2, the European Commission began enforcing new AI Act transparency obligations: chatbots must disclose they're AI, deepfakes must be labeled, and AI-generated content must carry machine-readable marks. The rules also require disclosure when AI is used for emotion recognition or biometric categorization. Penalties for non-compliance can reach €15 million or 3% of global turnover, though content made before August 2 is exempt and systems already on the market get until December 2 to comply.
3. DeepSeek V4 Flash 0731 Exits Preview, Beats Its Own Pro Model
DeepSeek's 284B-parameter mixture-of-experts model V4 Flash officially left preview status this week with a re-post-trained checkpoint tuned for agent workloads. It scored 82.7 on Terminal-Bench 2.1 — 14.7% ahead of the V4-Pro preview — plus 76.7 on Cybergym and 68.7 on DSBench-FullStack. Pricing holds at $0.14/M input and $0.28/M output tokens, undercutting most frontier-lab rivals by a wide margin.
4. Claude Sonnet 5's Intro Pricing Ends September 1
Anthropic's introductory rate for Claude Sonnet 5 — $2/M input, $10/M output tokens — expires August 31, jumping 50% to $3/M and $15/M on September 1. Compounding the increase, Sonnet 5's new tokenizer counts up to 35% more tokens for the same text than its predecessor, meaning workloads that look cheap today could land above old baselines within weeks even with a sticker price that reads "unchanged."
5. OpenAI Field Report: Coding Agents Deliver Up to 60x Speedups in Scientific Software
A joint OpenAI–academic report documents eight real deployments where coding agents (mostly Codex, some paired with Claude Code) modernized neglected research codebases. One genomics QC tool, RustQC, dropped runtime from over 15 hours to under 15 minutes — a 60x speedup — while another pipeline ran nearly 99x faster on its core compute step. The report cautions that agents executed well but couldn't judge scientific correctness on their own; human researchers still had to define and verify what "correct" meant.
6. Synthetic-User Startup Simile Raises $200M at $2B Valuation
Simile, which builds AI foundation models simulating human behavior for market research, closed a $200 million Series B at a $2 billion valuation — just five months after a $100 million Series A. The round was led by Greenoaks with Index Ventures returning; the company says revenue is up fivefold since its public launch, running tens of millions of behavioral simulations for Fortune 100 clients.
// KEY TAKEAWAYS
Frontier labs are racing on two fronts at once — raw capability (Astra's math proofs, DeepSeek's agent benchmarks) and cost/regulation pressure (Sonnet 5's price hike, the EU's transparency mandate). Meanwhile investors keep betting big on AI-native infrastructure and simulation plays like Simile, even as OpenAI's own coding-agent study underlines that human judgment, not speed, is still the bottleneck on trust.