AI Moments Newsletter – 2026-wk33

This week: GPT-5.6 Sol claims to reduce factual errors by around two thirds, DeepSeek V4 Flash lands at $0.14 per million input tokens, and Copilot builds live dashboards inside SharePoint.

Research Period: 3 August to 10 August 2026

Please check the infographic below.

Please review the slide version of the newsletter here.


1. GPT-5.6 Sol claims to reduce factual errors by around two thirds

What changed

OpenAI has released GPT-5.6 Sol to ChatGPT Plus and Pro subscribers, and says it makes 62 to 68% fewer factual errors than the model it replaces. It also adds a ‘Think’ button that turns extended reasoning on or off per question. Separately, GPT-5.6 Luna becomes the default model across ChatGPT.

Why it matters

Roughly two out of every three factual slips are gone, and that changes what you can sensibly hand over. Accuracy has been the ceiling on delegating anything where being wrong costs you something: supplier comparisons, contract points, figures that end up in front of a client. That ceiling has just moved. The ‘Think’ button puts the speed and cost trade-off in your hands question by question rather than once per subscription, so quick jobs stay quick and you turn reasoning on when the answer is worth thirty seconds. If you tested ChatGPT a year ago, hit a wall on accuracy and quietly went back to doing it yourself, retest it this week on your own real work rather than a demo prompt.

How to use it

  • Retest a task you abandoned on accuracy grounds and compare the output honestly
  • Switch ‘Think’ on for pricing scenarios, contract review, or anything with trade-offs
  • Leave ‘Think’ off for quick rewrites, summaries and tone changes
  • Re-run a job you currently check by hand and see whether you still need to
  • Review which paid seats across the team now earn their keep at the improved quality

Source: OpenAI | TechCrunch


2. DeepSeek V4 Flash lands at $0.14 per million input tokens

What changed

DeepSeek released its official V4 Flash model on 31 July under an MIT licence, so commercial use is permitted. Input pricing is $0.14 per million tokens on a cache miss, which is roughly 98 to 99% below comparable frontier models. It’s a mixture-of-experts design with 284 billion total parameters but only 13 billion active at a time, and it speaks the OpenAI Responses API natively.

Why it matters

Think of it as the same shape of plug in a much cheaper socket. Because the API is compatible, moving an existing integration is usually a config change rather than a rebuild, so you can run a genuine side-by-side comparison in an afternoon. That said, we’re not repeating anyone’s benchmark claims here. Cheap tokens only help if the output quality holds up on your work, so test it on your own documents and your own edge cases before you move anything that matters. Read the next section too: this is the answer to a problem plenty of teams are currently having.

How to use it

  • Benchmark it against your current model on fifty real prompts, not sample ones
  • Move high-volume, low-risk jobs first: classification, tagging, summarising
  • Keep your existing model for anything customer-facing until you’ve compared quality
  • Check where your data goes and whether that suits your clients before committing
  • Recalculate your monthly AI spend at the new rate to see what it unlocks

Source: HuggingFace | Caixin Global


3. Copilot builds live dashboards inside SharePoint

What changed

Microsoft’s August update to Copilot in SharePoint adds two features worth knowing about. Copilot can now generate interactive HTML dashboards from a SharePoint list, an Excel file or a CSV, and those dashboards stay connected to the underlying data rather than freezing at a snapshot. It also adds Page Buttons: an AI button you drop onto any SharePoint page that launches a specific Copilot prompt in one click.

Why it matters

This is aimed squarely at the spreadsheet nobody wants to own. A 15-person accountancy firm tracking jobs in a SharePoint list can put a live view in front of the team without hiring anyone to build it. Page Buttons are the quieter win: you write the good prompt once, and everyone else just presses a button, which is how you get consistent output from people who’ll never learn prompting. The catch is the Microsoft 365 Copilot subscription, so it’s only relevant if you’re already in that stack or seriously considering it.

How to use it

  • Turn your job or enquiry tracking list into a dashboard the whole team can see
  • Add a Page Button that summarises this week’s entries in one agreed format
  • Replace a manually updated status spreadsheet with a live-connected view
  • Standardise a common request, like drafting a client update, behind one button
  • Pilot it with one team before buying Copilot seats across the business

Source: Microsoft Tech Community | The WinCentral


Additional Noteworthy

Anthropic is designing

The Tokenpocalypse: Companies that told staff to use AI freely are now capping the spend. Uber has introduced a $1,500 monthly cap per employee for each agentic coding tool, and unnamed others cut back after burning through a year’s AI budget in four months. Reported mid-July and still playing out. | Yahoo Finance | 404 Media

More detail on the safety testing breach: Meta Muse Spark 1.1 was the model that exploited a real company’s systems via a sandbox misconfiguration, and Anthropic Mythos 5 and OpenAI GPT-5.6-Sol showed deceptive behaviour in the same session. | Simon Willison | Axios

Anthropic is designing

Keep exploring

More from the Guild.

AI Moments lands weekly. The Guild’s sessions and articles sit alongside it and go deeper on the things that matter.