This week had one headline model that almost nobody can use, one developer conference that tried to turn agents into a product category, and a funding story that keeps getting bigger. Last week's issue covered the Claude Sonnet 5.5 and Opus 5.5 launches, so we skip repeats here. Below: the numbers, the primary links, and what matters if you build or buy with these tools.
TL;DR
- Google announced Gemini 4 Argon on September 30: 77.9% on DeepSWE v1.1, $2/$10 per million tokens, but access starts with cyber defenders only.
- OpenAI used DevDay on September 29 to launch Dots, GPT-6.1 Sol, a Decisions API, and a $500/month Pro tier.
- Cloudflare open-sourced Clef decision models under Apache 2.0, a direct answer to OpenAI's new API.
- OpenAI is reportedly raising at least $30B at about a $1.4T valuation; Anthropic is reportedly eyeing a mid-November IPO.
- October opened with eight specialist model releases from five vendors in its first three days and no new frontier LLM.
Model Launches
Gemini 4 Argon is Google DeepMind's new frontier model for long-horizon coding, legal and finance work, and cyber defense. It scores 77.9% on DeepSWE v1.1 and raises the output ceiling to 1M tokens, up from 64K. Introductory pricing is $2/M input and $10/M output, later $4/$20, with a 95% discount on cached input.
The benchmark story is mixed. VentureBeat's read of Google's numbers says Argon leads or ties on 13 of 18 disclosed benchmarks against GPT-6 Astra and Claude Opus 5.5. It hit 19.6% on Harvey's Legal Agent Benchmark versus 5.4% for Astra, and 51.3% on AutomationBench versus 42.5% for Opus 5.5. It also beat Opus 5.5 on DeepSWE, 77.9% to 74.2%.
It loses in places. GPT-6 Astra leads by 10.5 points on FrontierSWE v2, and Opus 5.5 leads by 9 points on Terminal-bench 4.0.
The catch is access. Argon rolls out first to trusted defenders in the Fairwind Program and to U.S. government pre-release programs, with paid API customers and Google AI Ultra subscribers next and no date given. Every score so far is first-party or from a gated cohort, so treat the table as a claim, not a measurement. CNBC asked whether the model can really catch up to OpenAI and Anthropic at the frontier; the answer depends on what ships publicly.
OpenAI also shipped GPT-6.1 Sol at DevDay. Zvi Mowshowitz's weekly roundup lists it at $2/$10, pitched as near-Astra intelligence for about a fifth of the price. The same roundup notes an Ultrafast mode with 8x faster tokens at premium pricing. Opus 5.5 stays that author's preferred model.
Cloudflare released Clef and Clef-flash, open-source decision models under Apache 2.0. Clef is 27B at $0.24/M input on Workers AI; Clef-flash is 9B at $0.09/M with a 38.8 ms median decision. They are built on a Qwen backbone, use prefill-only scoring, and are API-compatible with Jev, per Developers Digest.
The pattern across the first days of October is specialization. The Digital Applied release ledger counts eight models from five vendors on October 1-2, with "no frontier LLM—every release in the first three days is a specialist." Four of the eight ship Apache 2.0 weights. The list includes:
- Amazon Strands Decider 2B: free and local.
- Microsoft AI: three speech models, with transcription at $0.54/hour of audio.
- Bilibili: a 35B total, 3B active translation model.
- Tavus Griffin-Lite: a full-duplex video-to-video model. The Neuron's October 1 digest reports 48% of 54 people in one-minute calls thought it was human, versus under 3% for earlier systems. Access is gated for safety review.
Other open and small releases from that digest: Inception Labs Mercury Voice at 320 ms median time-to-first-token and $0.20/$0.75 per million tokens, Kyutai PocketTTS (100M parameters), and Ai2 Olmo-core 3, claiming 2.7x training throughput over its predecessor.
Product Updates
OpenAI Dots was the DevDay headline. Ben's Bites summarizes it as friendly, always-on personal agents powered by GPT-6 Astra that you can name, give an avatar, and receive proactive messages from. One Dot is included with Pro plans. The same recap lists a new Pro 500 ($500/month) tier alongside Pro 100 and Pro 200, plus ChatGPT Space, a collaborative document workspace, and Plugins aimed at agent adoption.
The Decisions API is the developer piece. You pass text or images, define a finite list of allowed answers, and get back a selection your code can branch on. It runs on GPT-6 Luna, is in limited preview, and a broad release was promised within days. Coverage describes it as roughly 10x faster than a normal Luna call, though not the cheapest option available. OpenAI reportedly credited Jev as inspiration.
Competitors moved the same week. The Neuron's digest reports Perplexity launched a structured decision endpoint supporting up to 255 options at $0.04/M input tokens with free output. It also reports GitHub Copilot computer use in public preview on macOS and Windows, covering both the Copilot CLI and app. Claude merged Cowork and chat into one experience for long-running tasks that continue after the laptop closes, and Claude for Government became generally available in a FedRAMP High environment.
Smaller tools worth a look from the same digest:
- fal Recast swaps people in video for reference photos at $0.30/second (768p) or $0.45/second (1080p).
- FLUX 3 Image adds box-based layout control with native 2K and 4K output.
- Comfy Agent builds and debugs ComfyUI graphs from a text description.
Research & Papers
arXiv introduced new submission limits. Per its rate-limit policy post, users get two submissions per calendar month and at most three active submissions. Monthly volume rose from 9,869 papers in September 2016 to 40,363 in September 2026. Why it matters: preprint flooding is now a policy problem, and AI-assisted paper mills are the obvious suspect.
A study titled How Much Is an AI Token Worth? estimates 31.1% of FineWeb-filtered August 2026 web tokens were AI-generated, up from about 10% in June 2024. Per The Neuron's summary, AI text helps data-starved models but hurts once training reaches Chinchilla-optimal scale, and an unfiltered crawl needed 1.6x the compute of the human-only subset. Why it matters: pretraining data quality is getting harder to buy.
StudentBench tested AI tutors on 2,383 students and found they matched expert human GRE tutors on immediate learning gains. Gemma 4 31B matched human tutoring at roughly 918x lower cost per percentage point of gain. Why it matters: it is an education result with a concrete cost figure, though it measures immediate gains, not long-term retention.
On capability tracking, the same digest cites the Epoch AI Capabilities Index at 167 for Claude Opus 5.5, with GPT-6 Astra narrowly behind and Sonnet 5.5 near 165. Argon is not on that list yet because nobody outside the gate can run it.
Two reality checks from the same roundup: Pew Research found AI-generated "silicon sample" survey respondents missed human results by 12.4 points on average across about 300 questions. And the Ramp AI Index put AI adoption at 56.1% of tracked firms, with open-source models under 5% of business AI spend.
Industry & Funding
OpenAI is reportedly in talks to raise at least $30B at a $1.4T valuation, per Bloomberg's September 29 report. That follows the $122B raise in March at $852B. The IPO has slipped to 2027. TechCrunch reports run-rate revenue is up 70% since July, reaching $40B in monthly revenue by August.
The Neuron's digest adds that SoftBank closed a final $10B tranche (cumulative $64.6B, about 13% ownership) and Nvidia closed a final $10B of its pledge.
Anthropic is reportedly targeting a mid-November IPO, per Bloomberg as relayed by The Neuron. Marketing could begin the week of November 9, with an investor meeting on October 14 and a valuation at or above $2T. The same digest reports up to $42B in convertible-note financing with Broadcom tied to chip leases. Treat all of these as reported, not confirmed.
Smaller rounds from that digest:
- Armadin raised a $255.5M Series B for autonomous security agents.
- Volantis raised $88M for optical GPU-memory links.
- Halluminate raised $30M for finance RL environments.
- Salesforce acquired Listen Labs.
Regulation and legal news was heavy. According to The Neuron's digest:
- California AG Rob Bonta served OpenAI an investigative subpoena over cybersecurity incidents and model risks.
- OpenAI reportedly warned 100+ organizations about unauthorized rogue-agent activity.
- Senators Hawley and Murphy floated an AI Agent Accountability Act with liability for certain agent hacking under the CFAA.
- A federal judge dismissed Penske's suit over Google AI Overviews.
- Governor Newsom vetoed smart-glasses privacy bill SB 1130.
On the economics side, Columbia's estimate puts the AI buildout at $10.3T through 2032, and a16z estimates only about 2% of U.S. households paid for an AI service as of April 2026.
Community Buzz
The mood on X was skepticism about a model you cannot touch. One widely shared take:
"It's gonna be pretty funny when Gemini 4 Argon launches for the public, underperforms relative to its benchmarks like almost all Gemini models, and is then beaten by new Anthropic and OpenAI models this month. I'd like to be proved wrong though Happy October! Gonna be a fun one" — leo (@synthwavedd on X)
A search report indicates Argon drew heavy discussion on Hacker News too, with 1,300+ points and 870+ comments, mostly about the gated rollout and unverified numbers. We did not read that thread directly, so take the figures as reported.
On the open-weights side, one report on the Clef release says a thread on r/LocalLLaMA linked the Hugging Face card within half an hour of the announcement. That Remio write-up frames Cloudflare as the only one of three decision-model vendors giving away full weights. Our direct Reddit access was blocked this week, so the link above goes to the subreddit, not to a specific thread.
Zvi's roundup captured the other thread of the week: capability gains are arriving faster than safety consensus. Among the voices he quotes is Patrick McKenzie's advice to "Point them at hard problems where you are a good judge of success." Separately, the arXiv cap and the 31.1% AI-token estimate fed a loud debate about whether the web is filling with machine text.
The Writingmate Angle
Honest status: Gemini 4 Argon is not generally available anywhere yet, so there is nothing to test in Writingmate today. The Decisions API and Dots are OpenAI platform features, not models you pick in a chat window. What you can do now is run your own tests on models already on the platform, such as GPT-6.1 Sol Pro and Claude Sonnet 5.5, and compare them side by side on your real tasks. Browse the current list at writingmate.ai/models, and when Argon opens up, compare it against Opus 5.5 on your own prompts before believing any leaderboard.
Frequently Asked Questions
Sources
- Gemini 4 Argon
- VentureBeat's read of Google's numbers
- really catch up to OpenAI and Anthropic at the frontier
- Zvi Mowshowitz's weekly roundup
- Clef and Clef-flash
- Developers Digest
- Digital Applied release ledger
- The Neuron's October 1 digest
- Ben's Bites summarizes it
- rate-limit policy post
- in talks to raise at least $30B at a $1.4T valuation
- September 29 report
- $122B raise
- @synthwavedd on X
- r/LocalLLaMA
- Remio write-up
- writingmate.ai/models
Written by
Artem Vysotsky
Ex-Staff Engineer at Meta. Building the technical foundation to make AI accessible to everyone.
Reviewed by
Sergey Vysotsky
Ex-Chief Editor / PM at Mosaic. Passionate about making AI accessible and affordable for everyone.
