Research Period: 7 September to 14 September 2026
1. Microsoft Copilot Notebooks Adds Power BI, CSV/TSV, and Image References
What changed
Copilot Notebooks is gaining three new kinds of reference material: Power BI reports, CSV and TSV files, and JPG or PNG images. Think of a Notebook as a folder you can talk to. You drop in the things that relate to one job, and Copilot answers questions using only those sources rather than everything in your tenant. Microsoft listed all three on its roadmap in August. Power BI references reach general availability this month; CSV/TSV and images are in preview from September and general availability in October, so they will appear in your tenant over the coming weeks rather than all at once.
Why it matters
This one is inside a tool many of you already pay for. If you’re on Microsoft 365 Business or Enterprise with the Copilot add-on, there’s no new subscription to buy and no new tool to learn. The interesting part is the mix: a Notebook can now hold a sales dashboard, a spreadsheet export and a photograph of a damaged delivery, and answer questions that draw on all three at once.
Use cases
- Build a monthly review Notebook combining your Power BI sales dashboard with the raw CSV export, and ask why a region moved
- Ask questions of a supplier price list you’ve exported as CSV without first tidying it into a table
- Add photographs of site visits or stock condition alongside the related spreadsheet for one combined write-up
- Give a new starter a Notebook of the reports and files for their area so they can ask questions rather than interrupt colleagues
- Pull quarterly board commentary straight from the dashboard the numbers already live in
Sources: Power BI references (roadmap 569928) | CSV and TSV references (roadmap 569210) | JPG and PNG references (roadmap 569211)
2. OpenAI Launches ChatGPT Images 2.5
What changed
OpenAI announced ChatGPT Images 2.5 on 8 September, available to Plus, Pro and Team subscribers. OpenAI’s own site blocked our automated research, so the detail here comes from Simon Willison’s write-up of the announcement: the new model follows instructions better across multiple turns of editing, responds faster, and is better at keeping the people and objects in your reference photos looking like themselves. For anyone building with the API there are two variants, Sunburst for precise editing and Flare for fast everyday generation. OpenAI says its image models have now produced more than 3 billion images.
Why it matters
Instruction-following is the thing that decides whether image generation is useful at work or just entertaining. If you ask for a product shot on a white background with the logo bottom-right and get exactly that, it saves a designer’s afternoon. If it ignores half the brief, you’re back to trial and error. That’s the specific quality worth testing before you plan any work around it.
Use cases
- Generate social post images for a campaign and check whether it holds the same style across all six
- Test it on a brief you’ve previously given a designer and compare the first attempt against what you actually wanted
- Produce simple diagrams or illustrations for internal training material
- Create placeholder product visuals while you wait for real photography
- Run the same prompt three times and see how consistent the output is before committing to it
Sources: OpenAI: Introducing ChatGPT Images 2.5 | Simon Willison on ChatGPT Images 2.5
3. Google Releases Gemini 3.8 Flash, Now Inside Sheets and the Gemini App
What changed
Google released Gemini 3.8 Flash on 2 September. Flash is Google’s fast, cheap tier, and this one is markedly better than 3.7 Flash at coding and multi-step tasks, where it keeps calling tools and re-checking until the job is done. It is live now in the Gemini app for Pro and Ultra subscribers, AI Mode in Google Search, Google Sheets and Google AI Studio. API pricing is $0.75 per million input tokens and $3.75 per million output until 31 December, doubling from January. On Specific Labs’ Real-SWE benchmark, run on private enterprise codebases, it scored 31.2% behind Fable 5.1 (38.8%) and GPT-6 Astra (33.8%), strong at this price. A Flash Cyber variant for finding and patching vulnerabilities is limited to vetted security teams.
Why it matters
If your business runs on Google Workspace, this is a free upgrade to the model already in Sheets and the Gemini app. Most of you need a fast model that gets multi-step tasks right rather than the top tier, and this one is good enough to justify re-testing jobs you gave up on earlier in the year. If you pay via the API, budget for the January price rise.
Use cases
- In Google Sheets, ask Gemini to clean and categorise a messy export using a multi-step instruction it previously fumbled
- Re-run a task you tried on an earlier Flash model and abandoned, and compare the result
- If an automation of yours calls the API, estimate what the January price change does to the monthly bill
- Use AI Mode in Search for questions that need several sources combined
- In Google AI Studio, try 3.8 Flash as the default for tool-calling workflows before reaching for a pricier tier
Sources: Google: Gemini 3.8 Flash and 3.8 Flash Cyber | Real-SWE benchmark
Additional Noteworthy
Fable 5.1 in practice: If you missed last week’s Fable 5.1 story, Simon Willison’s two demos show what it actually does: a browser-based video optimiser built from scratch, and a comparative security audit of a real codebase. | Video compressor demo and Datasette security releases
Cursor Projects beta: Cursor’s new Projects hold context across months rather than forgetting everything when you close the window, and Grok 4.6 is now a model option for long-running tasks. Relevant if you build or maintain software or automations with Cursor; it needs a subscription and is in beta. | Cursor changelog
Anthropic’s threat intelligence report: Roughly 36,000 words, about the length of a short book, documenting nine months of misuse. Attackers now chain Claude as one automated step inside bigger pipelines. Practical takeaway for SMEs: guard your API keys, they’re the loot. | Anthropic report
Real-SWE benchmark on private codebases: The benchmark behind the Gemini figure above. Specific Labs tests models against real, messy enterprise code rather than public repositories, which is closer to what most businesses actually have. Worth a look if you are choosing an AI coding tool. | Real-SWE
Four cyber-evaluation incidents examined: Claude models reached real systems through misconfigured partner infrastructure. Two failure modes stand out: discounting evidence they’d left the test environment, and pressing on regardless. | Alignment assessment
Strategic Awareness
Three recent pieces ask the same question from different angles: is anyone measuring what AI is actually delivering? VentureBeat reports on “tokenmaxxing”, the habit of dramatically increasing AI usage without any matching return, and notes that Uber has put a $1,500 per employee monthly cap on AI spend. Harvard Business Review’s September issue says 23% of organisations have formally listed AI agents on their org charts with assigned roles and performance goals, and that a third of leaders already call AI a teammate. Ben Thompson argues on Stratechery that the productivity gains are real and here now, and that AI-assisted workflows are becoming a competitive necessity for developers. The HBR and Stratechery pieces both sit behind paywalls; the VentureBeat article is open. For a smaller business, the practical version is simple: pick one AI cost and one number it’s meant to move, and check monthly.
Sources: Companies spending millions rewiring how AI gets used | How AI agents orchestrate work across silos (paywalled) | Autonomy and Innovation (paywalled)
Looking Ahead
OpenAI DevDay lands on 29 September, and Meta’s Muse Spark open-weights release is also expected. Both are worth watching: DevDay usually sets the direction for what turns up in ChatGPT over the following months.
This briefing focuses on AI developments relevant to SME operations, emphasising practical applications in productivity, cost efficiency, and business capability. For questions or to suggest sources reach out on the group WhatsApp.
