Research Period: 31 August to 7 September 2026
Please review the slide version of the newsletter here.
1. Anthropic launches Fable 5.1 with 75% reduction in cache-read costs
What changed
Anthropic made Claude Fable 5.1 generally available on 1 September 2026, alongside Mythos 5.1, which is restricted to cybersecurity and life-sciences use cases. Cache-read pricing has been cut by 75%, from $1.00 to $0.25 per million tokens, while standard input and output pricing stays at $10 and $50 per million tokens. Cache writes are priced at $12.50 per million tokens for five-minute caches and $20 for one-hour caches.
Why it matters
Caching is how you avoid resending the same context, such as a long system prompt or a reference document, on every API request. The price cut makes that technique considerably more affordable for any business running high-volume document processing, customer support, or automated analysis through the API. Fable 5.1 also lifts its Terminal-Bench-Science score from 24.7 to 52.6, a substantial improvement in scientific reasoning, and it is the new default model for Claude Pro and Max subscribers at no extra cost.
Use cases
- Recalculate your API costs for any workflow that resends the same reference documents on every call
- Move document-heavy processing (contracts, specifications, policy manuals) onto cached prompts
- Build customer support automation with long system prompts that were previously too costly to run at volume
- Retest tasks that Claude previously handled poorly, particularly technical and scientific reasoning
- Review Mythos 5.1 if you work in life sciences or security, where its use-case restrictions apply
Sources: VentureBeat | Anthropic
2. OpenAI releases GPT-6 Astra, a model that operates software interfaces directly
What changed
OpenAI released GPT-6 Astra on 3 September 2026, first to trusted partners, then to paid users the same day. The headline capability is computer use: the model navigates interfaces, fills in forms, and moves through spreadsheets and web pages on its own. It is on ChatGPT Plus, Pro, Business and Enterprise plans, the OpenAI API, and AWS. API pricing is $10 and $50 per million tokens for input and output, 2.5 times the previous flagship rate. A restricted version initially declines certain cybersecurity-related prompts. OpenAI president Greg Brockman called it the start of artificial general intelligence.
Why it matters
Most AI tools still need you to paste the data in and copy the answer out. A model that drives the software itself removes that step, which matters for the repetitive work that sits between systems: rekeying figures from one application into another, filling in supplier portals, reconciling spreadsheets. This is the first time that capability has arrived inside a subscription most businesses already hold. Treat Brockman’s AGI framing as vendor positioning rather than a description of what the model does today. Computer use needs setting up and supervising before it saves anyone time, and volume work needs costing before you commit.
Use cases
- Pick one repetitive between-systems task (rekeying orders, updating a portal) and test whether Astra can complete it end to end
- Check whether your existing ChatGPT plan already includes it before adding any new spend
- Run any computer-use task on test data first, watching the steps rather than only the result
- Cost an API-based workflow at $10/$50 per million tokens before committing to volume
- Keep a human check on anything touching payments or client records
Sources: Fortune | Axios | CNBC
3. Perplexity Hybrid Compute processes sensitive files on your Mac, so nothing leaves the device
What changed
Perplexity released Hybrid Compute on 1 September 2026. It pairs cloud AI with a local model running on the Mac itself, and an on-device classifier called PII-Tracer, open-sourced on Hugging Face, decides what may go to the cloud and what must stay local. Sensitive files never reach Perplexity’s servers; only general search queries are sent. It requires an Apple Silicon chip, macOS 15 or later, a minimum of 24GB unified memory, and a Perplexity Pro subscription.
Why it matters
The most common objection to putting business documents through an AI tool is what happens to the data afterwards. Hybrid Compute answers that directly: HR records, client contracts, and financial data are processed on the machine, with only general web research going out. The 24GB memory requirement means this is for higher-specification Macs, though those are increasingly standard in SME settings.
Use cases
- Check your Mac’s specification against the requirements before planning any rollout
- Analyse HR documents such as contracts and appraisals without the content leaving the machine
- Review draft client proposals containing sensitive commercial terms on-device
- Process internal financial reports using Perplexity’s research capabilities locally
- Test on a low-stakes internal document first to confirm the on-device routing behaves as described
Sources: Perplexity Blog
Additional Noteworthy
Claude Code September update (Anthropic): Adds Remote Control, which lets external applications trigger and manage Claude Code sessions programmatically, plus VS Code integration improvements, larger inline editing limits, and new policy and skill diagnostics. Fable 5.1 is now the default Fable model in Claude Code. Automatic update for existing subscribers; Code features require Claude Pro or higher. | Claude Code changelog
Anthropic Enterprise Frontier Safeguards: Activity logs from your Claude usage stored in your own AWS, GCP, or Azure account under your encryption keys, with no Anthropic staff access to content; requires the Anthropic Enterprise tier. | Anthropic Blog
Google retires Google Assistant, Gemini replaces it: Google Assistant is being switched off on Android smartphones, tablets, Wear OS devices, and Android Auto globally, with no option to revert. | The Decoder
Ethan Mollick: The Twilight Factory (featured last week): If you missed it, Mollick’s piece on where humans belong in agentic workflows is worth your time. | One Useful Thing
Andrew Ng on coding agents: Writing in The Batch on 4 September, Ng sets out steering coding agents as one of the core AI engineering skills, and breaks it into five parts: directing the workflow, deciding how much autonomy to allow, reviewing what comes back, customising the environment, and understanding how agents fail. Relevant for any member using Claude Code or Cursor. | The Batch
Strategic Awareness
A Harvard Business Review study by Adam Peruta and Carrie Riby, run across 3,000 US consumers using controlled comparisons, found that AI-generated advertisements consistently underperformed human-created ads on purchase intent, brand recall, and emotional resonance. Notably, consumers could not reliably tell which was which, so the gap is invisible to the audience but shows up in the results. The authors treat this as a current capability gap rather than a permanent one. If you use AI for marketing content, human creative direction still matters for performance. Separately, Microsoft has restructured its reporting from September 2026 into a new “Agents and Infra” segment combining Azure, Microsoft 365, and GitHub, signalling that it sees agents as the primary delivery route for its cloud and productivity tools. Nothing to act on yet, but it suggests the M365 Copilot roadmap will accelerate; worth monitoring.
Sources: Harvard Business Review (paywalled) | MediaPost
Looking Ahead
Computer use is emerging as a common theme across vendors this week, so expect more announcements of this kind in the coming weeks. The on-device privacy story that Perplexity has opened is also likely to be followed by others.
This briefing focuses on AI developments relevant to SME operations, emphasising practical applications in productivity, cost efficiency, and business capability. For questions or to suggest sources reach out on the group WhatsApp.
