AI news daily brief: 2026-09-20
Seven stories today: two Anthropic developments (recursive self-improvement metrics placing 80 percent of production code under AI authorship, and a Bay Area wet lab), one security incident (Gemini's misconfigured sandbox breach), two product updates (ChatGPT in Microsoft Word for all plans, a $40M benchmarking round for Vals AI), one tooling change (GitHub Copilot model retirement), and one regulatory announcement (Trump's proposed AI Force).
Anthropic Reports 80 Percent of Its Production Code Now Authored by AI, With 30,000 Agents Running R&D Daily
The Anthropic Institute published "When AI builds itself" on September 18, 2026, documenting the company's progress toward recursive self-improvement, per the Anthropic Institute post. Claude-powered agents now author more than 80 percent of all code merged into Anthropic's production codebase, with engineers shipping roughly 8 times as much code per quarter as the 2021-to-2025 baseline. In a documented experiment, Claude Sonnet 5 improved an early version of Claude Opus 4.8 by testing over 50 ideas and producing a training method using 2,400 examples that reduced sycophancy, deception, and jailbreak susceptibility. The agent task horizon has been doubling every four months, reaching 12-hour tasks as of Claude Opus 4.6. Roughly 30,000 agents carry out daily research and engineering work at Anthropic as of August 2026. The company said full recursive self-improvement has not been achieved, but the metrics show a sustained exponential curve.
Full story: Anthropic recursive improvement metrics and 80 percent AI coding milestone
Google Confirms Gemini Hacked Three Real Companies During Authorized Cybersecurity Test
Google confirmed on September 19, 2026, that Gemini accessed the protected systems of three real companies during a cybersecurity evaluation conducted by Irregular, an Israeli security testing firm, per TechCrunch. A configuration error exposed Gemini to the public internet rather than an isolated sandbox. In one case Gemini guessed passwords until it gained access; in two others it found credentials in a public repository. Gemini stopped in each case after detecting it was operating on real production systems. Google notified federal authorities. This makes Google the fourth major AI lab after OpenAI, Anthropic, and Meta to disclose an Irregular-linked evaluation incident within a six-week window. No public evidence of in-the-wild exploitation has emerged. Whether Irregular publishes a full technical report and whether the incidents prompt industry-wide changes to evaluation sandbox standards are the two open questions.
Full story: Gemini hacked three companies during misconfigured cybersecurity evaluation
Anthropic Opens a Bay Area Wet Lab for Physical Biology Experiments and Launches a Life Sciences Verification Program
Anthropic is operating a wet biology lab in the Bay Area, per TechCrunch, confirmed by head of life sciences Eric Kauderer-Abrams on September 18, 2026. Claude models and robots help automate experimental workflows inside the lab. The focus is fundamental biology including rare diseases, not drug discovery. Alongside the lab confirmation, Anthropic launched a Life Sciences Verification Program giving vetted researchers access to Fable 5.1 and Mythos 5.1 for biological research tasks. Claude models are integrated with robotic lab equipment to accelerate experimental design and analysis. Anthropic has not named specific research partners. Whether the Verification Program expands beyond the current vetted cohort and whether the lab takes on external collaborators are the next developments to watch.
Full story: Anthropic Bay Area wet lab and Life Sciences Verification Program launch
OpenAI Brings ChatGPT to Microsoft Word for All Plans Including Free, with Enterprise Getting GPT-5.6 Sol at No Extra Cost Through September 30
OpenAI launched a ChatGPT sidebar add-in for Microsoft Word on September 17, 2026, available across all ChatGPT plans including Free, per RuntimeWire. The add-in handles drafting from notes, tightening arguments, formatting, and summarizing. Enterprise and Edu customers get GPT-5.6 Sol within Word at no extra cost through September 30, excluded from token limits. October 1 makes Word access the default for those tiers. The same plugin already powers Excel (May 2026) and PowerPoint (July 2026), completing the Microsoft Office suite. Whether OpenAI extends the free GPT-5.6 Sol trial past September 30, and whether Google responds with a comparable Docs integration for Gemini Ultra, are the two next milestones.
Full story: ChatGPT Word add-in available for all plans including Free
Vals AI Raises $40M from Andreessen Horowitz to Build a Confidential Benchmarking Standard for AI Models
Vals AI announced a $40 million Series A led by Andreessen Horowitz on September 19, 2026, at a $400 million valuation, per TechCrunch. Vals builds benchmarks that test model performance on real-world professional tasks across law, finance, coding, cybersecurity, and biosecurity, keeping test materials confidential to prevent models from being trained against them. Revenue grew eightfold relative to all of 2025. Headcount went from 8 to 25 employees during the first half of 2026. All five major AI labs now cite Vals results in their official model cards. The funding gives Vals runway to expand into higher-stakes evaluation domains including recursive self-improvement metrics, mental health applications, and the law of armed conflict.
Full story: Vals AI $40M Series A led by a16z for confidential AI benchmarking
GitHub Copilot Retires Six Models on October 19, Including GPT-5.5, Grok 4.5, and Gemini 3.7 Flash
GitHub announced on September 18, 2026, that six models will be deprecated across Copilot Chat, inline edits, agent mode, and code completions effective October 19, 2026, per the GitHub Changelog. Replacements are: Gemini 3.7 Flash to Gemini 3.8 Flash; GPT-5.5 and GPT-5.4 to GPT-5.6 Sol; GPT-5.4 mini and GPT-5 mini to GPT-5.6 Luna; Grok 4.5 to Grok 4.6. Enterprise and Business customers with default model enablement get alternatives applied automatically. Administrators who disabled the global default need to manually enable alternatives before October 19 to avoid a gap in model availability. No second deprecation wave has been announced.
Trump Announces an AI Force Modeled on Space Force and Vows to Name an AI Czar
US President Donald Trump posted on Truth Social on September 19, 2026, announcing he would create an "AI Force" modeled on Space Force to lead US AI development and keep the country ahead of China, per TechCrunch. Trump also announced he would name an AI czar and ran a poll on renaming artificial intelligence. No details on budget, staffing, or departmental placement were provided. Trump stated the new body would not hinder industry growth. Whether Trump follows through with an executive order naming the AI czar and defining the AI Force structure, budget, and authority is the next development to watch.
Sources
- When AI builds itself: Anthropic Institute, September 18, 2026
- Anthropic says its model Claude is helping to build the next version of itself: Washington Times, September 17, 2026
- Google's Gemini is the latest AI model to hack other companies: TechCrunch, September 19, 2026
- Google confirms Gemini hacked into three companies during cybersecurity test: 9to5Google, September 19, 2026
- Anthropic is operating a lab that conducts biology experiments: TechCrunch, September 18, 2026
- OpenAI adds ChatGPT to Word and opens its Office add-in to free users: RuntimeWire, September 2026
- ChatGPT Enterprise and Edu release notes: OpenAI Help Center, September 18, 2026
- Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking: TechCrunch, September 19, 2026
- Investing in Vals: Andreessen Horowitz, August 2026
- Upcoming deprecation of selected GitHub Copilot models in mid-October: GitHub Changelog, September 18, 2026
- Trump says it's time to rebrand AI with a new name and he's also creating an AI Force: TechCrunch, September 19, 2026
- Trump wants a new AI czar and an AI Force modeled on Space Force: Axios, September 19, 2026
