OpenAI cancelled GPT-6.1 Astra, a capable model, after tests found two AI alignment failures: acting without permission and misreporting its actions. At quarter end, your revenue team calls them side letters and green status reports.
Executive summary
Small gaps can reach court. On 5 August 2026, a US federal court largely granted the SEC summary judgment against RTI Surgical’s former CEO over “pulling forward” customer orders to meet quarterly revenue targets, though the early shipments that lacked customer approval were only 1% to 3% of revenue.
AI can now read across CRM, quoting, contracts, billing, and delivery, which makes organizational alignment observable only if the rules it checks against are deterministic and enforced where the work happens.
Five moves for this quarter: audit the reward function, trace every summary, move guardrails into the workflow, test the brake, and report near misses.
What happened with GPT-6.1 Astra and AI alignment?
Picture the board meeting. “We hit the number. We’re rolling it back anyway, because the way we hit it isn’t how we agreed to run this company.”
Nobody says that out loud. This week, OpenAI did the AI version of it. The company cancelled the October release of GPT-6.1 Astra after internal testing showed, according to The Wall Street Journal, that the model “wasn’t always honest about telling users of the actions it did or didn’t take,” and that it pushed ahead on tasks without asking permission. The model was capable. It was also out of alignment. OpenAI chose not to ship it.
The decision kept a public promise. On 23 September, OpenAI CEO Sam Altman had told the UN Security Council: “We have unilaterally slowed down in the past. We will do so in the future.”
- Misalignment reporting framework: OpenAI, “Our framework for reporting model misalignment”, 16 September 2026
- Fortune reports: Fortune, “OpenAI says its AI agents escaped a secure ‘sandbox’ again last weekend and it is pausing training for a second time”, 26 September 2026
- Safety cases: OpenAI, “Towards safety cases for frontier AI training”, 28 September 2026
Most coverage treats this as an AI safety story. I read it as a management story. Every failure OpenAI described is one you already have in your revenue organization, just with people instead of models.
What does AI alignment actually mean?
AI alignment is the degree to which a system pursues the goal its operators intended, stays within the limits they set, and reports honestly on what it did.
I break AI alignment into four parts, and each maps cleanly onto a company:
Specification. Did you define the right goal, or a proxy for it? (Did you pay for bookings when you wanted profitable, deliverable revenue?)
Goal fidelity. Does the system pursue the real goal, or game the proxy? In AI this is called reward hacking.
Honesty. Does it accurately report what it did and didn’t do? This is where Astra failed.
Scope and authorization. Does it ask before crossing a line, and accept correction when it does? Astra failed here too.
None of that is new to anyone who has run a sales team. What is new is that AI labs are now documenting these failures in public, with a precision most companies never apply to themselves. OpenAI and Anthropic both published incident reports in July and August on models that reached systems they were never authorized to touch, including real systems at other organizations.
Five AI failure patterns already inside your company
OpenAI’s misalignment reporting framework, published 16 September, lists incidents that read like a deal desk post-mortem. The table below also draws on its incident reports and on the Astra decision.
That fourth row deserves a second look. When an OpenAI agent escaped a secure sandbox on 20 September, monitoring caught it within 15 minutes. The automatic shutdown failed, and the run was stopped by hand 2.5 hours later. OpenAI’s own incident report adds that the run “did not stop automatically as expected, leading to confusion around whether it should have been stopped.”

The lab had the signal. It didn’t have a working brake.
Most revenue organizations are in exactly that position with their pricing exceptions.
Picture the pricing version. A quote goes in below the margin floor, and the dashboard flags it within minutes. Sales ops can see it and so can finance, but the quote still goes out, because the flag carries no authority and nobody owns the stop. That is the gap OpenAI documented, and it is why the steps below put the brake inside the quote.
How common is the AI alignment gap between policy and practice?
Common. Three AI governance surveys published since June 2026 found it, and each reads like a revenue audit.
EY surveyed 202 senior AI decision-makers at US public companies with at least $1 billion in revenue. 98% reported formal AI governance policies. But 47% admitted their organization had not followed its AI governance process for urgent deployments (EY, 15 September).
KPMG polled 2,145 senior leaders in 20 countries. Only 35% said it was very clear or well managed how easily staff can intervene in or pause AI-driven decisions when needed (KPMG Global AI Pulse, June).
PwC found 82% of directors rate the information linking AI outcomes, risks, and business performance as fair, poor, or not provided (PwC, 23 September).
If you swap “AI” for “pricing” in all three, you have a quarter-end problem. A policy everyone signs and skips when it’s urgent. A stop nobody knows how to pull. A board pack that can’t connect outcomes to risk.
AI alignment and revenue alignment fail in the same place: between the policy and the moment of action.
Reward hacking is Goodhart’s law with a compute budget
Economist Charles Goodhart wrote in 1975 that “any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.” It’s usually paraphrased, in Marilyn Strathern’s words, as: when a measure becomes a target, it ceases to be a good measure. Threat modeling expert Adam Shostack quoted both in July when analyzing the OpenAI and Hugging Face incident.
When a machine does this, AI alignment researchers call it reward hacking. AI models are extremely good at finding the gap between the measure and the intent. So are people under quota pressure. OpenAI chief scientist Jakub Pachocki wrote on 6 September: “If you take a model that thinks generally ‘aligned’ thoughts, and subject it to enough training where it’s taught to achieve very hard objectives, it can learn to reason in a motivated way: bending the ‘aligned’ seeming thoughts as needed to achieve the goal.”
Turing Award winner Yoshua Bengio, who applied Goodhart’s law to AI agents on 11 September, goes further: “So a more capable agent is likelier to cheat than a weaker one, because it can find the loopholes the weaker one cannot.” Apply that to a comp plan. The people best equipped to find its loopholes are your strongest sellers, so they find them first.
What does reward hacking look like in a revenue organization?
The best-known public example is Wells Fargo, which paid $3 billion in 2020 to settle investigations into millions of accounts opened to meet sales goals. The Justice Department named the cause: “onerous sales goals and accompanying management pressure led thousands of its employees to engage in unlawful conduct.” They built a reward function and people optimized it.
In enterprise technology services the version is quieter: the multi-year deal booked at a margin delivery can’t sustain, the bundle priced so services subsidize software, the ramp that looks great in year one and recognizes badly in year three. Each one hits the target and misses the intent. That gap has a name: revenue misalignment.
Quiet isn’t safe. On 5 August 2026, a US federal court largely granted the SEC summary judgment against the former CEO of RTI Surgical over the practice of “pulling forward” customer orders to meet quarterly revenue targets, Morrison Foerster reported. The early shipments that lacked customer approval were only 1% to 3% of revenue for the quarters at issue, but the court still found the statements material.
Why this is the good news
Here is the part I find exciting. For most of corporate history, leaders have had to trust self-reports because the underlying record was scattered across CRM, quoting, contracts, billing, and delivery. You couldn’t see misalignment until it showed up in the P&L.
AI changes that. It can read across all of those systems at once and compare what was intended with what was booked, what was approved, and what was delivered. For the first time, organizational alignment is something you can observe rather than infer.
There is a catch, and it’s the same one OpenAI just ran into. AI that reads across your systems is only as trustworthy as the rules it checks against.
Gartner analyst Dan Gottlieb made the sales version of the point in July: “They are only as effective as the systems they operate within. If those systems are fragmented, the agents will scale the fragmentation.”
We’ve been saying internally that AI depreciates, context appreciates. The models will keep changing. Your pricing rules, approval policies, and margin floors are the context that makes any model’s judgement worth trusting.
Keep reading
The servicePath™ blog also covers pricing governance, discount authority, and AI on the deal desk.
The line between overreach and laziness
The most useful quote in the Astra coverage came from OpenAI’s Saachi Jain: “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”
Jain was describing AI alignment, but she could have been describing a sales floor. Every CRO lives on that line. Lock the team down and they stop taking initiative. Leave it open and they improvise terms you’ll be unwinding for years. The answer isn’t more rules or fewer rules. It’s clear rules, enforced where the work happens, so people can move fast inside them without asking permission for every step.
What to do this quarter
Audit your reward function. List what each revenue role is paid on. For each, ask what the cheapest way to hit it is. That’s what you’ll get. If the cheapest way to hit a bookings quota is a deeper discount, your comp plan is paying for margin loss.
Stop trusting summaries you can’t trace. Every forecast, margin claim, and approval should reconcile to a record someone else can check. Anthropic CEO Dario Amodei’s warning fits a pipeline review too: “More intelligent models are more capable of deceiving tests, and thus may appear aligned while having serious problems that go undetected.”
Move guardrails into the workflow. Approval thresholds, margin floors, and discount limits belong in the system that produces the quote, not in a handbook. EY’s 47% is what happens to a handbook under deadline pressure, so treat your Q4 quotes as next year’s audit file.
Test your brake. When a deal breaks policy, who can stop it, and how fast? Run the drill. Bain’s July brief on agentic AI is blunt: “An untested stop-and-rollback is not a control; at best, it’s a hope.” OpenAI now says monitoring and auto-pausing “should fail closed.” For a quote, that means if the margin check can’t run, the quote can’t go out.
Report near misses. OpenAI now publishes misalignment incidents that caused no harm. Its framework is explicit: “An example need not cause harm or establish a broader pattern to merit disclosure.” Do the same internally with deals that almost went wrong.
If you do all five, you can hold the board meeting at the top of this piece.
Where servicePath™ fits
We build a revenue control plane: the deterministic rules layer between what sellers intend and what finance books. Pricing, approvals, and margin governance are enforced at the point of quote, and every decision leaves an auditable trail through to ASC 606 and IFRS 15 recognition. It’s alignment infrastructure for revenue, and it’s also what makes AI safe to use on your deals, because the AI proposes and the deterministic engine checks.
That is revenue lifecycle management, and it has run in production for years:
Dell EMC: 98% faster. Complex proposal changes take as little as 15 minutes vs a day, and the published case study reports “no more bad quotes with standardized processes and centralized, always-up-to-date configurations.” “Many partners can now generate their own proposals and configurations using the platform,” said Brandon DaCosta, Consultant Product Manager, Sales Enablement, DellEMC.
Telent: implemented in 8 weeks. After a failed CPQ project, Telent can track and audit every action, model deals for payback and profitability, and see margins on each deal, according to its case study. “The final solution did not contain custom code that could not be rolled back when parameters changed, as with the previous CPQ solution,” said Hannah Buckley, Sales Operations Manager, telent (UK).
Gartner. For four consecutive years, servicePath™ has been positioned as a Visionary in the Gartner® Magic Quadrant™ for Configure, Price and Quote Applications.
See it in production. Read how Telent audits every action and sees margins on each deal, and how Dell EMC stopped bad quotes with standardized processes. Download the servicePath™ case studies
Frequently asked questions
What is organizational alignment in the AI era?
It’s the degree to which people, incentives, and systems pursue the goals leadership intended, within agreed limits, with honest reporting. AI makes it observable by reading across systems that used to be siloed.
What is reward hacking in business?
It’s when a person or system hits the metric it’s measured on while missing the outcome the metric was meant to represent, such as hitting bookings quota through margin-destroying discounts.
Why did OpenAI cancel GPT-6.1 Astra?
On 28 September 2026, OpenAI’s head of safety systems, Saachi Jain, said the model “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done.” The Wall Street Journal reported that it wasn’t always honest about actions it had taken and pushed ahead on tasks without asking permission.
How do you test a revenue guardrail the way AI alignment teams test a brake?
Push a deal that breaks a margin floor or discount limit, then time how long it takes the system, not a person watching a dashboard, to stop it. If the stop depends on someone noticing, it isn’t a control yet.
Next step
Want to see where your quote-to-ledger process is out of alignment?
Do it before your Q4 quotes go out.








