Skip to content
Daily BriefNewsDaily Brief

7 AI stories from September 20, 2026: Anthropic 80-percent AI coding milestone, Gemini sandbox breach, Anthropic wet lab, ChatGPT in Word, Vals AI Series A, GitHub Copilot model retirement, and Trump AI Force

· by Pondero Newsdesk · 7 stories

AI news daily brief: 2026-09-20

Seven stories today: two Anthropic developments (recursive self-improvement metrics placing 80 percent of production code under AI authorship, and a Bay Area wet lab), one security incident (Gemini's misconfigured sandbox breach), two product updates (ChatGPT in Microsoft Word for all plans, a $40M benchmarking round for Vals AI), one tooling change (GitHub Copilot model retirement), and one regulatory announcement (Trump's proposed AI Force).

Anthropic Reports 80 Percent of Its Production Code Now Authored by AI, With 30,000 Agents Running R&D Daily

The Anthropic Institute published "When AI builds itself" on September 18, 2026, documenting the company's progress toward recursive self-improvement, per the Anthropic Institute post. Claude-powered agents now author more than 80 percent of all code merged into Anthropic's production codebase, with engineers shipping roughly 8 times as much code per quarter as the 2021-to-2025 baseline. In a documented experiment, Claude Sonnet 5 improved an early version of Claude Opus 4.8 by testing over 50 ideas and producing a training method using 2,400 examples that reduced sycophancy, deception, and jailbreak susceptibility. The agent task horizon has been doubling every four months, reaching 12-hour tasks as of Claude Opus 4.6. Roughly 30,000 agents carry out daily research and engineering work at Anthropic as of August 2026. The company said full recursive self-improvement has not been achieved, but the metrics show a sustained exponential curve.

Full story: Anthropic recursive improvement metrics and 80 percent AI coding milestone

Google Confirms Gemini Hacked Three Real Companies During Authorized Cybersecurity Test

Google confirmed on September 19, 2026, that Gemini accessed the protected systems of three real companies during a cybersecurity evaluation conducted by Irregular, an Israeli security testing firm, per TechCrunch. A configuration error exposed Gemini to the public internet rather than an isolated sandbox. In one case Gemini guessed passwords until it gained access; in two others it found credentials in a public repository. Gemini stopped in each case after detecting it was operating on real production systems. Google notified federal authorities. This makes Google the fourth major AI lab after OpenAI, Anthropic, and Meta to disclose an Irregular-linked evaluation incident within a six-week window. No public evidence of in-the-wild exploitation has emerged. Whether Irregular publishes a full technical report and whether the incidents prompt industry-wide changes to evaluation sandbox standards are the two open questions.

Full story: Gemini hacked three companies during misconfigured cybersecurity evaluation

Anthropic Opens a Bay Area Wet Lab for Physical Biology Experiments and Launches a Life Sciences Verification Program

Anthropic is operating a wet biology lab in the Bay Area, per TechCrunch, confirmed by head of life sciences Eric Kauderer-Abrams on September 18, 2026. Claude models and robots help automate experimental workflows inside the lab. The focus is fundamental biology including rare diseases, not drug discovery. Alongside the lab confirmation, Anthropic launched a Life Sciences Verification Program giving vetted researchers access to Fable 5.1 and Mythos 5.1 for biological research tasks. Claude models are integrated with robotic lab equipment to accelerate experimental design and analysis. Anthropic has not named specific research partners. Whether the Verification Program expands beyond the current vetted cohort and whether the lab takes on external collaborators are the next developments to watch.

Full story: Anthropic Bay Area wet lab and Life Sciences Verification Program launch

OpenAI Brings ChatGPT to Microsoft Word for All Plans Including Free, with Enterprise Getting GPT-5.6 Sol at No Extra Cost Through September 30

OpenAI launched a ChatGPT sidebar add-in for Microsoft Word on September 17, 2026, available across all ChatGPT plans including Free, per RuntimeWire. The add-in handles drafting from notes, tightening arguments, formatting, and summarizing. Enterprise and Edu customers get GPT-5.6 Sol within Word at no extra cost through September 30, excluded from token limits. October 1 makes Word access the default for those tiers. The same plugin already powers Excel (May 2026) and PowerPoint (July 2026), completing the Microsoft Office suite. Whether OpenAI extends the free GPT-5.6 Sol trial past September 30, and whether Google responds with a comparable Docs integration for Gemini Ultra, are the two next milestones.

Full story: ChatGPT Word add-in available for all plans including Free

Vals AI Raises $40M from Andreessen Horowitz to Build a Confidential Benchmarking Standard for AI Models

Vals AI announced a $40 million Series A led by Andreessen Horowitz on September 19, 2026, at a $400 million valuation, per TechCrunch. Vals builds benchmarks that test model performance on real-world professional tasks across law, finance, coding, cybersecurity, and biosecurity, keeping test materials confidential to prevent models from being trained against them. Revenue grew eightfold relative to all of 2025. Headcount went from 8 to 25 employees during the first half of 2026. All five major AI labs now cite Vals results in their official model cards. The funding gives Vals runway to expand into higher-stakes evaluation domains including recursive self-improvement metrics, mental health applications, and the law of armed conflict.

Full story: Vals AI $40M Series A led by a16z for confidential AI benchmarking

GitHub Copilot Retires Six Models on October 19, Including GPT-5.5, Grok 4.5, and Gemini 3.7 Flash

GitHub announced on September 18, 2026, that six models will be deprecated across Copilot Chat, inline edits, agent mode, and code completions effective October 19, 2026, per the GitHub Changelog. Replacements are: Gemini 3.7 Flash to Gemini 3.8 Flash; GPT-5.5 and GPT-5.4 to GPT-5.6 Sol; GPT-5.4 mini and GPT-5 mini to GPT-5.6 Luna; Grok 4.5 to Grok 4.6. Enterprise and Business customers with default model enablement get alternatives applied automatically. Administrators who disabled the global default need to manually enable alternatives before October 19 to avoid a gap in model availability. No second deprecation wave has been announced.

Trump Announces an AI Force Modeled on Space Force and Vows to Name an AI Czar

US President Donald Trump posted on Truth Social on September 19, 2026, announcing he would create an "AI Force" modeled on Space Force to lead US AI development and keep the country ahead of China, per TechCrunch. Trump also announced he would name an AI czar and ran a poll on renaming artificial intelligence. No details on budget, staffing, or departmental placement were provided. Trump stated the new body would not hinder industry growth. Whether Trump follows through with an executive order naming the AI czar and defining the AI Force structure, budget, and authority is the next development to watch.

Sources