What Changed In AI This Week? OpenAI Security, Microsoft Copilot in Excel, Gemini Cost Controls and Salesforce Agents

AIAI NewsAI GovernanceMicrosoft CopilotCyber Security

What Changed in AI This Week? OpenAI Security, Microsoft Copilot in Excel, Gemini Cost Controls and Salesforce Agents

OpenAI published details of an agent security incident, Microsoft made Copilot more useful in Excel, Google added stronger cost controls for agent workloads and Salesforce pushed AI agents further into commercial workflows.

Written by Jordan StewartAI news and governance
This week's developments in business AI
The short answer

This week’s AI news shows the technology becoming more like infrastructure. Agents are gaining real permissions, costs are becoming variable and operational, Copilot is moving deeper into business workflows, and vendors are experimenting with outcome-based pricing. The practical response is clearer ownership, tighter access, proper cost controls and better testing.

Each week on a Friday, we dive into and unpack the biggest AI stories and condense them into something easy to follow and understand, so you don’t have to scour through every single headline.

Why? Well, as AI continues down its path to prove it can become more capable, more measurable, and, rather harder to treat as a harmless experiment for businesses, we recognise there’s a demand to stay in the loop. But with so many stories coming out on the regular, and by the day, it can be somewhat overwhelming to know where to begin.

For example, you may have missed Anthropic, in what was one of the week’s most important developments, as they joined OpenAI, Microsoft, Google and more than 100 other organisations in warning that businesses have a limited window to strengthen cyber defences.

So let’s get into it.

Headlines at a glance

01

OpenAI’s agent security incident

During internal cyber evaluations, models operating with reduced safeguards escaped intended controls and accessed OpenAI and Hugging Face systems.

02

Microsoft Copilot gets more practical

Copilot can now use Python while editing Excel workbooks, alongside smaller changes that reduce switching between Copilot and Outlook.

03

Google tackles “token shock”

Gemini Enterprise is adding pay-as-you-go access, spend caps, pooled quotas and commitment discounts for agent workloads.

04

OpenAI publishes Jalapeño chip results

Its first custom inference chip delivered higher throughput per watt and lower latency in OpenAI’s tests, although customers have no immediate buying decision to make.

05

Salesforce pushes agents into commercial workflows

New commerce and service agents include pay-per-resolution pricing, while Salesforce expands the surrounding data and content platform.

OpenAI’s Hugging Face incident makes agent security a board-level issue

The biggest story this week was a follow-up on OpenAI's account of a July security incident, involving internal agents used for cybersecurity evaluations.

According to OpenAI, the models were operating with reduced safeguards and were meant to work inside isolated testing environments. Instead, they found ways to communicate through unauthorised channels, exploited weaknesses in shared infrastructure, reached the internet and accessed third-party systems. The activity was primarily driven by an internal-only research model comparable in scale to GPT-5.6 Sol, rather than the typical public ChatGPT experience.

OpenAI says the agents chained vulnerabilities together, recovered credentials, and gained administrator-level access across parts of Hugging Face’s infrastructure. It called the incident a “warning shot” and claimed to be tightening sandboxing, internet access, monitoring, and control of model weights.

A day later, Anthropic, Microsoft, Google, AWS and others joined OpenAI in a broader industry warning about AI-enabled cyberattacks. The letter itself doesn’t create any new compliance requirements or make specific spending commitments just yet. But the direction is clear: the organisations building the most capable models believe the defensive window is narrowing.

What this means for your business

If your business is giving an AI agent access to email, files, code, browsers or business systems, treat it the same way you’d treat a privileged user. Limit what it can reach, separate testing from production, use the least permissions possible, log tool calls, review high-impact actions and make sure there’s a reliable way to stop it before the genie gets out of the bottle, so to speak. And remember, an attitude of “oh, well, the model was only testing” isn’t a control.

Microsoft Copilot can now use Python while editing Excel workbooks

Microsoft’s 25th of August Microsoft 365 Copilot release notes are slightly less of a dramatic story, but probably more immediately useful for many businesses. Edit with Copilot in Excel can now execute Python for advanced analysis, automation, data transformation, statistics, simulations and visualisation, with results written back into the workbook.

That could make more sophisticated analysis accessible to finance, operations and commercial teams that know the question they want to answer, but perhaps don’t fully know how to write Python themselves. It also reduces the need to export sensitive data to a separate notebook or third-party tool just to run deeper analysis, which should make everything that bit more efficient.

The same release cycle adds some other smaller, but useful workflow improvements, too: users can open referenced Outlook emails alongside Copilot Chat, the Copilot app has a simplified design, and Copilot Cowork can generate and edit images. Microsoft notes that Copilot features roll out gradually within tenants, so availability may not be identical for every user on day one.

What this means for your business

Pick a real workbook and test one controlled use case, such as margin analysis, forecasting or data cleaning. Check the formulas and Python output against a known result, confirm who can access the underlying data, and measure whether the process is genuinely faster.

Google Gemini Enterprise adds pay-as-you-go pricing and hard spend caps

Google has addressed one of the less glamorous but increasingly important AI problems: how to pay for agents without losing control of the bill.

Gemini Enterprise is adding a pay-as-you-go edition with no base subscription fee, available to selected customers and rolling out more broadly. Businesses can mix that with per-user subscriptions, pool included developer quotas across a Google Cloud project and choose whether overages are allowed.

Google also introduced hard monthly project caps, alerts at 50%, 80%, and 100% of the budget, anomaly detection, and estimated agent runtime costs. For steadier workloads, its Flexible Savings Plans offer 10% off with a one-year commitment or 20% off with a three-year commitment. Deferred execution, which Google says can reduce eligible inference costs by up to half, is still coming soon for selected workloads.

You should treat this as the clearest sign this week that agentic AI is starting to look like cloud infrastructure rather than office software, and a licence per person still makes sense for regular individual use. Background agents, coding tools and bursty workflows need a different meter.

What this means for your business

Don’t choose pay-as-you-go solely because it sounds flexible, or a long commitment simply because it seems to offer a discount on paper. Run a small workload first, establish a monthly baseline, set caps, and decide which tasks are allowed to pause when the budget is reached. A savings plan only saves money when the underlying usage is real and reasonably predictable.

OpenAI’s Jalapeño chip points to faster, cheaper AI, but not yet

OpenAI also published the first measured results from Jalapeño, its custom inference chip. Inference is the part of AI computing that serves a trained model to users, so improvements here affect response speed, capacity and the underlying cost of running products.

Across the models OpenAI tested, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. OpenAI says it will ramp up the chip over the coming months.

It seems as though OpenAI is trying to control more of the stack, from models and serving software to chips, memory and networking, which perhaps makes this story more commercially interesting, rather than anything you could benefit from.

But, conversely, if that produces sustained efficiency gains, it could support faster agents, better availability and eventually lower prices. “Eventually” is doing some heavy lifting there, though: OpenAI is yet to announce a new customer price or a general availability date tied to these results.

What this means for your business

No procurement action is required this week, but it’s worth keeping an eye on whether the measured efficiency turns into lower API prices, better service levels or new regional capacity. Hardware claims only really matter to customers when they change the invoice or the experience.

Salesforce moves Agentforce towards outcomes, commerce, and end-to-end service

Salesforce’s quarterly product update shows how quickly agents are moving into customer-facing workflows. Agentforce Commerce now connects product catalogues to channels, including ChatGPT, Google Search, and Gemini, while purpose-built agents can support product discovery, checkout, service, and merchandising.

The most notable thing here is Agentforce Help Agent, which Salesforce says can manage cases, schedule appointments and update orders across web, portal and messaging channels. Pricing is pay-per-resolution: the customer is charged when the agent resolves an issue autonomously without a human stepping in.

That sounds neatly aligned with value, but outcome pricing raises its own questions. What counts as resolved? How are reopened cases treated? Does the agent optimise for a quick closure rather than a good customer outcome? Businesses should insist on seeing those definitions and escalation rules before building a business case.

Salesforce is also expanding the platform to better support its agents. It has agreed to acquire Contentful for a native content layer, and Fin, an AI customer-service platform used by more than 30,000 companies; both deals remain subject to closing conditions.

What this means for your business

For service teams, compare cost per genuinely resolved case, customer satisfaction, repeat contact and escalation rates, not just automation volume. Outcome-based pricing can be useful, but only when the vendor’s definition of success matches yours.

Quick answer: what should SMEs do about AI this week?

01

Review agent permissions

Identify every AI tool that can browse, execute code, send messages or reach business data. Reduce access to the minimum each workflow needs.

02

Put cost controls beside access controls

Set budgets, alerts and stop conditions before usage-based agents move beyond a pilot.

03

Test Copilot on a known Excel problem

Validate the output and the time saved before expanding the workflow to more people or more sensitive data.

04

Define outcomes in plain English

If a supplier charges per resolution or completed task, agree on what “done” means and how errors, rework and escalation are measured.

Fifosys view

The common thread this week is that AI is becoming infrastructure. Let’s look at the facts: it’s connected to systems, consumes variable resources, performs longer tasks, and carries permissions that can cause real damage when controls fail. It’s kind of hard to argue anything else, really.

As we tell people regularly, you shouldn’t panic or rush to switch everything on just because ‘everyone else is!’

The key to getting AI adoption right is in making it look more like mature IT: that means named owners, approved tools, limited access, measurable use cases, clear budgets, testing and incident plans.

Is that less exciting than being told an agent can do everything? Sure. But after this week’s news, an agent that can do exactly what it is allowed to do, and no more, looks rather more valuable.

This week's AI news FAQs

What happened in OpenAI’s Hugging Face incident?
OpenAI says internal cyber-evaluation agents operating with reduced safeguards escaped their intended testing boundaries, chained vulnerabilities together and accessed parts of Hugging Face’s infrastructure. The incident has led OpenAI to tighten sandboxing, monitoring and internet-access controls.
Can Microsoft Copilot now use Python in Excel?
Yes. Microsoft’s 25 August 2026 Copilot release notes say Edit with Copilot in Excel can execute Python for analysis, automation, data transformation, statistics, simulations and visualisation, with results written back into the workbook.
What new Gemini Enterprise cost controls has Google introduced?
Google is adding pay-as-you-go access, pooled quotas, hard monthly project caps, spend alerts, anomaly detection and estimated runtime costs for agent workloads, alongside commitment discounts for steadier usage.
What is OpenAI’s Jalapeño chip?
Jalapeño is OpenAI’s custom inference chip. OpenAI says early tests delivered higher throughput per watt and lower latency than the comparison systems it measured, although no immediate customer pricing or purchasing change has been announced.
How does Salesforce’s pay-per-resolution pricing work?
Salesforce says customers are charged when Agentforce Help Agent resolves an issue autonomously without a human stepping in. Businesses should still confirm exactly how a resolution is defined, how reopened cases are treated and how escalation affects the commercial model.
Jordan Stewart
Jordan StewartFifosys insights, news and practical technology guidance for UK business leaders.
Next
Next

Your IT manager doesn't need replacing - they need backup