OpenAI decided one model was too dangerous to release. Three days later it put a hacking model on sale.
Last issue ended with OpenAI saying something no AI company had said before: that it could not rule out "critical" hacking ability in Astra, its unreleased model, and was locking it in isolated machines until governments could test it.
On August 10 the same company launched GPT-5.6-Cyber. It is built for exactly the work Astra was locked up over: finding unknown holes in software and writing the code to exploit them.
The difference is who can buy it. It is sold only through a vetted programme called Daybreak Red, which requires identity checks, legal undertakings, approved-use limits, and from September 1 a physical security key to log in. OpenAI rates this one "high" capability rather than "critical," the tier below Astra's.
What it can do is not in dispute. Before launch, OpenAI says the model found two previously unknown flaws in the part of Chrome that runs the code on web pages, which chained together could break the browser's defences. One has a public identifier, CVE-2026-15903. It also found more than 400 privilege-escalation bugs in a widely used operating-system kernel, and three critical database flaws.
How often each model does the offensive-security work rather than refusing it. OpenAI's own evaluation, reported at launch · Source: OpenAI, Aug 10 2026
The clearest measure of what changed is that test: how many advanced security tasks a model will actually complete rather than refuse. The ordinary version of its flagship completes 1.5%. This one completes 95%.
The flaws it found were in the browser on your machine, and they are now fixed. That is the honest case for building this. The equally honest worry is that a tool this good at finding holes is only ever one leaked login away from the people who use holes. OpenAI's answer is the vetting, which is why the identity checks and the September security-key requirement matter more than the model does. Watch the gate, not the gadget.
GPT-5.6-Cyber is a variant of GPT-5.6 Sol trained for offensive security work, launched Aug 10 2026 and available only through Daybreak Red. Daybreak, OpenAI's cyber-defence programme, now splits into two tiers: Blue gives approved defenders frontier general-purpose models with system-level cyber guardrails removed; Red is the only route to the purpose-trained model. Both require identity verification, approved-use restrictions and legal attestations, with hardware security keys mandatory for all individual Daybreak accounts from Sept 1 2026.
On OpenAI's internal Advanced Cybersecurity Completion Rate evaluation: Sol with safeguards 1.5%, Daybreak Blue 2.0%, GPT-5.5-Cyber 57.3%, GPT-5.6-Cyber 95.0%. Under the Preparedness Framework it is classified High, not Critical: OpenAI states it "improved over GPT-5.6 Sol on some specialized cyber tasks... but not sufficiently to reach our Critical threshold." Plain Sol still beats it on vulnerability-report quality, on ExploitBench at the 300-turn setting, and on token efficiency.
Reported findings include two V8 vulnerabilities (V8 is Chrome's JavaScript engine), one assigned CVE-2026-15903, at least five in an unnamed mobile operating system, three critical database vulnerabilities allowing remote code execution, and more than 400 kernel privilege-escalation vulnerabilities. No per-token price is published. Disclosure: Human Terms is written with help from Claude, made by Anthropic, a competitor of OpenAI, which appears here and again in Follow the Money.
A free AI good enough to matter now runs on a laptop.
On August 14 Alibaba released Qwen3.8-27B. Anyone can download it, it costs nothing, and it runs on a decent recent laptop with about 24GB of memory. The download is roughly 17GB, about four films. It handles images and video as well as text, and on the public tests it scores near models that cost real money.
32 to 48GB is the comfortable range. Benchmark claims are vendor-reported and not independently replicated · Source: model card and independent hardware write-ups, Aug 2026
This is the version of the price collapse that reaches you directly. Every AI you use today sends your words to a company's computers. This one runs on your machine, so the work never leaves the building. For a small business handling client files, or anyone who would rather not post their private life to a server, that is a genuinely different offer. It also means "the good AI" is no longer something only rich companies can hold.
Qwen3.8-27B is a dense 27-billion-parameter model under the Apache 2.0 licence with native vision-language input and a 262K context window, about 17GB quantised at Q4_K_M. Practical local requirements are 24GB of unified memory or VRAM as a floor, 32 to 48GB as a comfortable range. Reported scores include 61.7% on SWE-Bench Pro, 90.3% on LiveCodeBench v6 and 89.2% on GPQA Diamond.
Musk's Grok pulled level with OpenAI's flagship Grok 4.6
Grok 4.6 arrived on August 12 and scored 61 on Artificial Analysis, an independent scoreboard that rates models across many tests. That ties OpenAI's top model and sits two points behind Anthropic's Claude Opus 5, which has led since late July. It costs $2 per million words of input against OpenAI's $5.
Four labs are now within two points of each other. When we started this newsletter one company was clearly ahead; today the honest answer to "which AI is best" is that it barely matters, and price and habit decide it. That is what a market looks like when the product becomes a commodity.
Grok 4.6 launched Aug 12 with a 500,000-token context window at $2/$6 per million input/output tokens below 200K prompt tokens, doubling to $4/$12 for requests above that threshold, applied to all tokens in the request. Its 61 on the Artificial Analysis Intelligence Index matches GPT-5.6 Sol; Claude Opus 5 leads at 63.
Google shipped another middleweight while its flagship stayed missing Gemini 3.7 Flash
On August 13 Google released Gemini 3.7 Flash, a fast, cheap model for everyday work. The same day, a Forbes column noted that Gemini 3.5 Pro, the flagship meant to compete at the top, is now roughly five months past its original target with no confirmed date. Last issue we passed on a rumour that it would arrive August 12. It did not.
We keep reporting this because the pattern is now the story. Google has the most users, the most data and its own chips, and it still cannot ship the model that would prove it leads. Being enormous does not settle this contest, which is why the other three keep catching up.
Gemini 3.7 Flash shipped Aug 13 at introductory pricing through Dec 31 2026. Gemini 3.5 Pro remains in limited preview on Vertex AI against an original June 2026 target. Benchmark figures throughout this section are vendor-reported or third-party aggregations rather than independent replications.
Oracle is cutting staff again. Not because AI did their jobs, because AI has to be paid for.
On August 12, reporting indicated Oracle is preparing another round of job cuts, with some teams facing double-digit reductions. Managers were told to identify who goes before September 1, the start of the company's second fiscal quarter.
The reason sits in the accounts rather than in any AI system. In the year to May, Oracle spent $55.7 billion building data centres, up from $21.2 billion the year before, close to a tripling. It spent $23.7 billion more cash than it generated, raised $43 billion in debt, and expects to raise roughly $40 billion more. It has already cut about 21,000 jobs in the prior fiscal year, a 13% fall in headcount.
Fiscal year to May 31 2026. About 21,000 jobs already went that year; a further round was reported for before Sept 1. Oracle declined to comment · Source: Oracle results as reported Aug 12 2026
This is not a struggling company. Revenue grew 17%, and its cloud infrastructure business grew 77%. The shares are down about 26% this year anyway, because investors are watching the borrowing.
Most AI job stories describe a machine doing someone's work. This is a different mechanism and probably a more common one: payroll is being cut to help pay for the machines. The tell is that it lands in a growing company, and that the deadline is a quarterly reporting date rather than the arrival of any particular technology. If your employer is spending heavily on AI infrastructure, the number worth watching is not what the AI can do. It is how the spending is being financed.
The planned cuts were reported Aug 12 2026; Oracle declined to comment. Fiscal 2026 ended May 31 2026 with capital expenditure of $55.7bn against $21.2bn in fiscal 2025, cash outflow exceeding generation by $23.7bn, $43bn of debt raised, $5bn of stock sold, and roughly $40bn of further borrowing or equity issuance anticipated. Headcount fell about 21,000 to roughly 141,000, and the company attributed those earlier reductions to "the adoption and deployment of AI technologies." Restructuring charges were $1.8bn with a total anticipated up to $2.1bn. Revenue grew 17% and cloud infrastructure 77%. Shares are down about 26% year to date. This is one company's reporting, not a sector-wide statistic.
The first proper trial of an AI well-being app found it helped loneliness and did nothing for depression.
Published August 13 in NEJM AI, a journal from the New England Journal of Medicine group. Researchers gave 486 undergraduates at three American universities either access to an AI app called Flourish or their normal support, then followed them for six weeks.
The students told to use it twice a week reported real gains: more resilience, a stronger sense of belonging, feeling closer to their community, and less loneliness. They also held steady on measures where the control group slipped. On depression, anxiety and stress, there was no difference at all.
486 students, six weeks, randomised and preregistered. Two authors are employees and directors of the company that makes the app; the senior author holds shares. Not a treatment study, and not medical advice · Source: Cachia et al., NEJM AI, Aug 13 2026
One more thing belongs in the plain read rather than the footnotes. Two of the authors are employees and board members of the company that makes the app, and a third is an adviser who holds shares. The researchers who ran recruitment at the three campuses declare no such interests, and the trial was preregistered, meaning the questions were locked in publicly before the data came in. Both facts are worth holding at once.
If you have a teenager using one of these, this is the most useful finding available. On the ordinary human stuff, feeling less alone, feeling part of something, there is now evidence an app can help. On clinical conditions, there is none, and this trial was not designed to find any. An app that makes a lonely week better is a good thing. It is not treatment, and the study's own authors do not claim it is. Nothing here is medical advice.
Cachia, Zhao, Hunter, Wu, Lin and De Freitas, "AI for Proactive Mental Health: A Multi-Institutional, Longitudinal, Randomized Controlled Trial," NEJM AI, published Aug 13 2026 (DOI 10.1056/AIoa2501293). Preregistered, six weeks, 486 undergraduates at Foothill College, Chapman University and the University of Washington, randomised to app access or care-as-usual control.
The intervention arm reported significantly greater positive affect, resilience and social well-being (belonging, closeness to community, reduced loneliness) and was buffered against declines in mindfulness and flourishing. No significant condition-by-time effects were observed for depression, anxiety or stress in this general-population sample. Competing interests: two authors are employees and directors of Flourish Science Inc., the public benefit company that built the app, and the senior author advises the company and holds shares; the site investigators who led recruitment and data collection declare none. The trial ran in autumn 2024, so the app tested is nearly two years older than the version on sale today. A general-population student sample is not a clinical population, and the trial was not powered or designed as a treatment study.
Should anyone be allowed to sell a hacking machine?
This week's lead did not land quietly. OpenAI's argument is that defenders are already losing: attacks using AI are multiplying, the people breaking in are not waiting for permission, and security teams at banks, hospitals and utilities need tools of the same class. The Chrome flaws its model found are the exhibit. Those holes existed whether or not anyone looked, and now they are patched.
The objection is that a company which spent July explaining how its own models broke into other people's systems has spent August selling that ability as a product, with a price list. Vetting is only as good as the vetting, and every gated system in history has eventually admitted someone it should not have. There is also a structural complaint: OpenAI now profits from both sides of the same problem.
Both sides accept the same underlying fact, which is unusual and clarifying. Nobody argues the capability can be un-invented. The disagreement is only about whether keeping it in a small number of vetted hands is safer than the alternative, and the alternative is not "nobody has it." It is that the people willing to break the law get there anyway, a little later.
This argument decides something concrete for you, which is whether the software you rely on gets fixed faster than it gets broken. There is no neutral option here and no way to opt out. The thing to watch is whether the vetting holds, because the entire case for doing it this way rests on that one control, and the first time it fails the argument is over.
The dual-use tension is explicit in OpenAI's own framing, which describes Daybreak Blue as the recommended starting point for most defenders and Red as the route for authorised vulnerability research and exploit validation. Reported customers for OpenAI's security offerings include established security firms. Neither Anthropic, Google nor Meta currently sells a comparable purpose-trained offensive security model. This item deliberately shares no figures with the Big Picture item above: the capability numbers are reported there and the argument here.
The two biggest AI companies are about to ask the public to buy their shares, and they look nothing alike.
OpenAI is moving toward a stock market listing as early as September at a valuation above $1 trillion, with Goldman Sachs and Morgan Stanley leading. Its filing shows roughly $2 billion a month in revenue, about $25 billion a year, against a projected loss of around $14 billion for 2026. It does not expect to generate more cash than it spends until 2030.
Anthropic is the mirror image. Backers are now discussing a valuation of $2 trillion or more, against $965 billion when it last raised money in May, with a listing expected in October. Its case rests on a single quarter: the company told investors it expected $10.9 billion of revenue for the April-to-June quarter and $559 million of operating profit, its first.
Anthropic's $559m quarterly operating profit is a projection given to investors in May, not an audited result. Not investment advice · Sources: reporting on OpenAI's confidential filing (Aug 17) and on Anthropic's backers (Aug 14)
Three cautions on that profit figure, because it is doing a lot of work. It was a projection given to investors in May rather than an audited result. It is operating profit, which comes before several kinds of cost, not the bottom line. And critics pointed out at the time that a temporarily discounted deal for computing power flattered the quarter, with the company itself noting it might not stay profitable across the full year as spending rises.
For scale: a $1 trillion company would be worth about as much as the fifth largest company on the American stock market. Two AI firms are proposing to arrive at roughly that size within weeks of each other.
If you have a pension or an index fund, you are about to own a piece of both, without being asked. It is worth knowing they are running opposite experiments. One is betting that spending faster than you earn buys a position nobody can take from you. The other is betting that showing a profit, even a thin and contested one, is what turns an expensive story into a company. Within a year or two the market will tell you which was right, and the answer will move the value of a lot of ordinary retirement accounts. Nothing here is investment advice.
OpenAI filed a confidential S-1 in June 2026; reporting on Aug 17 describes a public debut targeted as early as September 2026 with a valuation range from about $852bn to above $1tn, roughly $2bn of monthly revenue, more than 900 million weekly ChatGPT users, enterprise contracts above 40% of revenue, a projected 2026 loss near $14bn, and no positive cash flow expected before 2030.
Anthropic's $10.9bn revenue and $559m adjusted operating profit are projections it gave investors, first reported by CNBC on May 20 2026, implying a margin near 5.1%, against $4.8bn of revenue in the prior quarter; compute spending per revenue dollar reportedly fell from 71 to 56 cents. Ed Zitron's May 21 2026 critique argues the figure is non-GAAP, excludes costs, and depends on a discounted compute arrangement. Its Series H in May 2026 valued the company at $965bn. Reported run-rate figures vary widely by source and measure and are not asserted here. Disclosure: Human Terms is written with help from Claude, made by Anthropic, one of the two companies in this item.
California decided which AI bills live. The ones about children advanced; the one about artists died.
On August 13 California's two money committees held what the legislature calls suspense hearings: a day when bills with a cost attached are either sent onward or quietly killed. Of 29 live AI bills, 24 moved forward and five were held.
"Held in committee" is worth understanding once, because it is where most bills actually die. There is no vote against the bill and no debate. It simply is not called, and the session ends.
Among the five held: AB 412, which would have required disclosure of copyrighted training data · Source: Assembly and Senate appropriations suspense results, Aug 13 2026, published as unofficial
The bills that survived cluster around children. AB 2023 and SB 1119 set safety rules for chatbots used by minors, SB 300 targets sexual content in chatbots, and SB 867 would keep them out of toys. The most notable death is AB 412, which would have required AI companies to document the copyrighted material used to train their systems. It has now failed twice.
Everything that survived has to clear a floor vote before the session ends at midnight on August 31. Meanwhile in Washington, the request we covered last issue, asking the Speaker to make the AI chiefs testify under oath, has gone unanswered.
Two things about your life were decided in a Sacramento committee room in one afternoon. If you have a child using a chatbot, the rules governing that are moving. If you write, draw, record or photograph for a living, the bill that would have told you whether your work is inside these systems is dead again. Notice which one had organised opposition. And notice that the state legislature moved on both in an afternoon while Congress has not scheduled a hearing.
The Assembly and Senate appropriations committees held simultaneous suspense hearings on Aug 13 2026. Of 29 active AI bills, five were held in committee: AB 412 (documentation of copyrighted training data), AB 2545 (AI worker impact assessments), SB 1015 (deepfake-induced extortion), SB 1146 (AI in health-related advertising) and SB 1181 (AI and youth mental health). Advancing measures include AB 2023 (do pass, 6-1) and SB 1119 on chatbot safety for minors, SB 300 on preventing chatbot sexual content (do pass, 13-0), SB 867 on chatbots in toys (do pass as amended, 11-0) and SB 813 on AI standards. The session ends at midnight on Aug 31 2026. Suspense-file results are published as unofficial at the time of the hearing. Bills that pass both chambers still require the governor's signature, a separate step not covered here.
Three things worth keeping an eye on.
- SeptemberOpenAI's listing window opens. When a company actually goes public its filing stops being confidential, so for the first time outsiders see audited numbers rather than briefed ones. That single document will settle several arguments in this issue.
- Aug 31California's floor votes, before midnight. Twenty-four surviving AI bills have about two weeks. What passes there tends to become the national default, because companies rarely build one product for California and another for everyone else.
- OpenWhether any other lab starts selling offensive security tools. Anthropic, Google and Meta have all stayed out so far. If one follows OpenAI, this stops being one company's judgement call and becomes an industry norm, which is the point at which regulators usually arrive.
This week a company sold a tool that finds security holes, arguing the defenders need it more than the attackers do. When you hear about a technology that cuts both ways, what would actually reassure you? Hit reply to the email. I read every one.