In early September 2026 OpenAI said GPT-6 Astra is the first model to reach the Critical level of cybersecurity capability under its own Preparedness Framework. Microsoft made the model available for business work in Foundry the same week. Most coverage focused on what the model can break. For an owner buying automation, the useful question is narrower: what is this agent allowed to touch in your business, and who approves the steps that matter?

What the Critical label means, and who applied it

OpenAI’s deployment safety hub states the classification directly: Astra is the “first model to reach the Critical level of cybersecurity capability under our Preparedness Framework.” The framework’s Critical threshold covers work carried through without a person guiding each step.

Two qualifications belong next to that sentence. The classification is the maker’s own, not an independent grade. And OpenAI says advanced offensive cyber work is gated rather than open in normal production. Both are vendor statements. Read them as the vendor’s risk signal and check them against your own situation.

The part that affects a working business

Astra is described as strong at computer use: reading what is on a screen, filling forms, updating records, moving between applications that have no clean connection to each other. Microsoft’s Foundry announcement makes the same point for business work — open-ended goals turned into multi-step work across apps.

That is the capability most owners actually want. It is also the capability that needs limits. An agent that can click and type can act on the wrong record, and it can be steered by text it reads on a page it visits. OpenAI reports better resistance to that kind of injected instruction than in the previous model, and still does not call the problem solved.

One finding deserves attention from anyone planning to supervise an agent by reading its reasoning: OpenAI reports that monitorability has decreased relative to GPT-5. Supervision that depends on the model explaining itself honestly is weaker ground than supervision that gates actions.

The controls are the product

Both vendors publish their control lists. Those lists are the part worth reading.

OpenAI describes stricter isolation, monitoring of full task trajectories, checks run before internal use, and production monitoring that can stop unauthorised activity. Microsoft describes scoped credentials, approved resources, human checkpoints for consequential actions, activity records matched to risk, plus identity, role-based access and private networking.

Microsoft also states plainly that these controls help you configure safeguards and do not remove your responsibility to choose the right ones for your situation. That sentence is the honest summary of the whole topic.

Five questions to ask any automation vendor

  1. Least privilege. Which systems and records can the agent reach, and which are out of scope? Ask for the list, not a reassurance.
  2. Human checkpoints. Which actions require a person? Money movement, customer identity changes, bulk deletions and anything sent to a customer belong on that list.
  3. Records. Can you reconstruct what the agent saw, wrote and changed, and stop a run in progress?
  4. Untrusted input. How does the system treat text on a web page, an inbound email or an attachment — as information, or as instructions?
  5. Refusals and approvals. What is refused by default, and who approves an exception?

An owner with those five answers in writing can compare two vendors properly. An owner without them is buying on tone.

What this changes for a small operator

Not much this week, and quite a lot over the next year. Nobody running a plumbing, HVAC or service business needs a frontier model to send an estimate reminder. The work that pays off now is still ordinary: one process, written down, with the steps a person must approve marked as such.

That groundwork is exactly what makes stronger agents safe to adopt later. A business that already knows which actions need a human has a place to put a more capable model. A business that does not will hand it everything.

Sources

Kush AI Automation builds this way by default: scoped access, drafts a person approves, and a record of every step. See what we build, or book a free call and bring one process you want handled.