ChatGPT Business
Paid users
10
Premium / heavy users
2
Cursor Teams
Paid users
5
Premium / heavy users
1
ChatGPT billing
ChatGPT: 8 Standard + 2 Premium. Cursor: 4 Standard + 1 Premium. Estimate excludes tax, usage overages, and API purchases.
Complete analytical reference
The full business case behind the presentation: economics, ChatGPT plan differentiation, current model benchmarks, coding-harness mechanics, governance, and a measurable 10-week rollout.
Decision and economics
Recommendation, interactive-seat assumptions, alternatives, and decision boundaries.The step beyond free chat is not only “better answers.” It is a managed operating environment: dependable capacity, frontier-model access, company context, enforceable data policy, usage visibility, and one bill.
| Layer | Buy | Why now | Who gets it |
|---|---|---|---|
| General AI | ChatGPT Business | Company knowledge, higher limits, secure workspace, admin controls | Knowledge workers who already use AI chat |
| Software delivery | Cursor Teams | AI-native agent harness, broad frontier models, repo context, team policy | Developers and technical builders |
ChatGPT Business
Paid users
10
Premium / heavy users
2
Cursor Teams
Paid users
5
Premium / heavy users
1
ChatGPT billing
ChatGPT: 8 Standard + 2 Premium. Cursor: 4 Standard + 1 Premium. Estimate excludes tax, usage overages, and API purchases.
Standard: $25/user/month. Premium: $125/user/month with 5× usage and no five-hour usage limit.
Standard and Premium seats can coexist in one workspace and be assigned or reassigned. This directly supports a “baseline + power-user” policy.
Official seat documentationStandard: $40/user/month. Premium: $120/user/month with 5× the Standard Agent allowance.
Admins can upgrade or downgrade individual users. Upgrades are immediate and prorated; included usage remains per-user rather than pooled on Teams.
Official team pricing| If this is the priority | Best starting point | Reason |
|---|---|---|
| Fastest path beyond free chat | ChatGPT Business | Small migration burden; immediate gains in capacity, company context, and governance |
| Best integrated coding-agent experience | Cursor Teams | AI-native editor, multi-file agent loop, model routing, shared rules and skills |
| Lowest-friction GitHub / multi-IDE rollout | Copilot Business | Native GitHub administration and broad IDE reach |
| Strongest terminal-native steering and isolation | Claude Code Team | Rich subagents, hooks, permissions, sandboxing, and a tightly co-designed Claude harness |
| Maximum model freedom or local/BYOK | Kilo or VS Code BYOK | Hundreds of providers/models, local endpoints, and direct provider billing |
ChatGPT business case
Free-versus-Business comparison, GPT-5.6 model evidence, paid capabilities, and governance.The strongest argument is operational, not cosmetic. Business turns fragmented personal use into an environment the company can fund, govern, connect, and measure.
| Dimension | Free personal workspace | ChatGPT Business | Case to make |
|---|---|---|---|
| Usage and continuity | Limited messages, uploads, deep research, memory, context, Codex, and slower image generation | Higher included limits; Standard or 5× Premium seats; workspace credits can extend usage | Fewer interruptions during real work |
| Model access and quality | GPT-5.6 Luna is the current default; no access to Sol reasoning tiers | GPT-5.6 Sol at Medium, High and Extra High, plus GPT-5.6 Sol Pro; Luna/Terra/Sol in Work and Codex | Harder work can use stronger reasoning, not only a faster chatbot |
| Company knowledge | Personal connections and ad hoc file uploads; no shared company workspace policy | Company knowledge across enabled apps such as Slack, SharePoint, Drive, GitHub, HubSpot, and Asana, with citations | Answers can be grounded in current internal sources |
| Connector governance | Configured by each individual | Admins manage plugins/apps and permissions; access respects each user's source-system permissions | Connectors become governable rather than shadow integrations |
| Data handling | Personal-workspace training sharing is enabled by default, with user opt-out | Workspace inputs and outputs are excluded from model training by default | A safer default for company information |
| Administration | No organization-level roles, usage visibility, or policy | Owners/admins, centralized billing, usage analytics, spend controls, MFA and SAML SSO | One accountable owner and policy surface |
| Right-sizing | One free tier per person | Mix and reassign Standard ($20 annual / $25 monthly) and Premium ($100 annual / $125 monthly) | Pay for intensity where it is proven |
| Model | Positioning | Where users encounter it | Quality evidence |
|---|---|---|---|
| GPT-5.5 Instant | Fast conversational baseline and previous default | Replaced GPT-5.3 Instant for all users in May 2026; now being superseded by the 5.6 rollout | 52.5% fewer hallucinated claims than GPT-5.3 Instant on OpenAI's high-stakes internal prompts; no direct same-harness 5.6 comparison published |
| GPT-5.6 Luna | Fastest, lowest-cost 5.6 model | Default and Think for Free/Go; Work and Codex on eligible paid plans | Good high-volume default; lower peak reasoning than Terra and Sol |
| GPT-5.6 Terra | Balanced capability, speed, and cost | Work and Codex; not selectable in ordinary ChatGPT conversations | Near-Sol coding performance in OpenAI's published model-level evaluations |
| GPT-5.6 Sol | Flagship for complex reasoning, research, coding, and long-running work | Instant with automatic reasoning plus Medium, High and Extra High on Business | Best 5.6 family result across the three engineering benchmarks below |
| GPT-5.6 Sol Pro | Highest-capability ChatGPT option for difficult, longer-running workflows | Pro picker option on Business, Enterprise, and individual Pro | No clean public ChatGPT-specific apples-to-apples score located; do not substitute the API's Sol Ultra result |
Higher is better. GPT-5.5 here means the reasoning/Thinking model, not GPT-5.5 Instant. These are model-level evaluations reported by OpenAI, not a guarantee of ChatGPT product results.
| Evaluation | GPT-5.5 Thinking | GPT-5.6 Luna | GPT-5.6 Terra | GPT-5.6 Sol |
|---|---|---|---|---|
| AA Intelligence Index v4.1 · index | 54.8 | 51.2 | 55.0 | 58.9 |
| AA Coding Agent Index v1.1 · index | 76.4 | 74.6 | 77.4 | 80.0 |
| SWE-Bench Pro · resolved | 59.4% | 62.7% | 63.4% | 64.6% |
| DeepSWE v1.1 · resolved | 67.0% | 67.2% | 69.6% | 72.7% |
| Terminal-Bench 2.1 · pass | 85.6% | 84.7% | 87.4% | 88.8% |
| BrowseComp · solved | 84.4% | 83.3% | 87.5% | 90.4% |
| OSWorld 2.0 · success | 47.5% | 45.6% | 50.2% | 62.6% |
Source: OpenAI GPT-5.6 launch report · June 2026. Results depend on reasoning effort and harness configuration.
| Capability | Free | ChatGPT Business | Procurement meaning |
|---|---|---|---|
| Deep research | Available with a small plan-dependent allowance | Expanded use; editable research plan, files, web, selected sites, and enabled apps with cited reports | A real research workflow, but capacity—not mere availability—is the paid distinction |
| Agent mode | Unavailable | 40 initiated messages per month before flexible overage, subject to current limits | Delegates browser-and-tool workflows rather than only producing an answer |
| Projects | Available; 5 files per project | Shared team projects, up to 40 files per project, and workspace privacy defaults | Business adds collaboration and governed data handling |
| Scheduled tasks | Limited cadence and flexible delivery windows; up to 3 active | Higher active-task limit, hourly/exact schedules, and eligible event triggers in Work; current help text conflicts between 10 and 15 active tasks | Useful for monitoring and recurring work, not just reminders |
| Create custom GPTs | Cannot create or publish new GPTs on personal plans | Members can create managed workspace GPTs when admins permit | Reusable internal assistants become company-managed assets |
| Company knowledge | No managed organization knowledge layer | Permission-aware answers across approved apps, with source citations; currently web-only | The clearest capability that is genuinely Business-specific |
| Workspace agents | Not available as a managed workspace capability | Create, preview, share, and schedule repeatable agents within the workspace | Moves from one-off prompts toward repeatable workflows |
| ChatGPT Work | Limited access | Multi-step agent for research, connected apps/files, and finished documents, sheets, presentations, reports, and Sites | Delegated deliverables rather than chat responses |
| Codex | Limited access; Terra where available | Expanded coding access with Sol, Terra, and Luna plus workspace controls | Useful secondary coding surface, though Cursor remains the dedicated editor proposal |
Call this “permission-aware company knowledge,” not a universal data warehouse. Quality depends on which apps are enabled, source freshness, indexing mode, and the user's existing access.
Move the conversation to Enterprise when the organization needs SCIM, EKM, domain verification, fine-grained RBAC, custom retention, invoice/PO terms, or more than the self-serve Business workspace limit.
“The team already uses free AI, so the choice is not adoption versus no adoption. It is unmanaged personal use versus a secure company workspace. ChatGPT Business gives us stronger models and dependable capacity, grounds answers in approved company sources, excludes workspace data from training by default, and consolidates billing and administration. Mixed Standard and Premium seats let us scale heavy users without overbuying for everyone.”
Model benchmarks
CursorBench, Artificial Analysis, OpenRouter caveats, cost, and model–harness interpretation.Coding quality is a property of the model–harness pair. CursorBench is useful for comparing models inside Cursor's own agent loop, while Artificial Analysis compares complete coding-agent systems. Neither replaces a matched pilot on your repositories.
Higher is better. CursorBench uses ambiguous, multi-file tasks derived from real Cursor sessions. It is first-party and not independently reproducible; effort levels and cost per task vary materially. Cursor warns that small gaps may not be statistically meaningful, so the 0.8-point spread across the top three is not decision-grade.
| Rank | Model | Best score | Interpretation |
|---|---|---|---|
| 1 | Grok 4.6 · Extra High | 70.8% | Current peak; $2.81/task in Cursor's published run |
| 2 | Claude Fable 5 · Max | 70.5% | Near peak, but $17.32/task |
| 3 | Claude Opus 5 · Max | 70.0% | Peak Claude value tier; $8.23/task |
| 4 | GPT-5.6 Sol · Max | 67.2% | Strong flagship result; $5.69/task |
| 5 | GPT-5.6 Terra · Max | 64.9% | Balanced 5.6 option |
| 6 | Gemini 3.7 Flash | 61.6% | High-throughput frontier option |
| 7 | Claude Sonnet 5 | 61.5% | Lower-cost Claude default |
| 8 | GPT-5.6 Luna | 61.1% | Fastest 5.6 tier |
| 9 | Kimi K3 | 60.8% | Open-weight, long-horizon option |
| 10 | GPT-5.5 | 58.4% | Previous OpenAI reasoning generation |
| 11 | Composer 2.5 | 56.1% | Lower peak score but $0.44/task and Cursor-native tuning |
Source: CursorBench 3.2 public leaderboard · snapshot checked 26 Aug 2026. The Grok 4.5 training-data caveat is omitted because this table uses Grok 4.6.
The top 3.6 percentage points span roughly $2.81 to $17.32 per task. Composer 2.5 scores 14.7 points below the leader but costs about one-sixth as much as Grok 4.6 Extra High and about one-fortieth as much as Fable 5 Max in Cursor's run.
CursorBench holds Cursor's harness broadly constant to expose model differences. It does not tell us how Cursor versus Claude Code versus Copilot performs with each product's native prompting, context management, safeguards, and verification loops.
Artificial Analysis Coding Agent Index v1.4: simple average of DeepSWE, Terminal-Bench 2.1, and SWE-Atlas-QnA pass@1. These rows compare complete systems, not isolated models.
| Agent harness | Representative model | Index | Cost / task | What it shows |
|---|---|---|---|---|
| Claude Code | Claude Opus 5 · xhigh | 68 | $8.17 | Highest of this current cross-check; strongest Terminal-Bench and repo-Q&A mix |
| Codex | GPT-5.6 Sol · max | 65 | $6.42 | Best DeepSWE result here and materially faster than the Claude Code run |
| Grok Build | Grok 4.5 · high | 64 | $2.44 | Near-frontier composite at much lower cost |
| Kimi Code CLI | Kimi K3 | 63 | $3.08 | Strong open-weight model–native harness pairing |
| OpenCode | Gemini 3.7 Flash · high | 60 | $1.27 | Strong terminal score and the lowest task cost in this set |
| Cursor CLI | GPT-5.5 · medium | 47 | $2.00 | An older Cursor/model configuration; absence of current Cursor models prevents a fair product ranking |
Coding harness assessment
Cursor, Copilot, VS Code BYOK, Kilo, and Claude Code across quality, steering, safety, and cost.Cursor, Copilot, VS Code BYOK, Kilo, and Claude Code increasingly overlap on frontier-model capability. The durable differentiator is how the harness builds the base prompt, gathers repository context, exposes tools, lets teams steer behavior, verifies results, recovers from failure, and gates risky actions.
Internal assessment, 1–5. It is a hypothesis for the pilot, not a vendor benchmark. Weights favor end-to-end agent work over autocomplete.
| Criterion | Weight | Cursor Teams | Copilot Business | VS Code + BYOK | Kilo Teams | Claude Code Team |
|---|---|---|---|---|---|---|
| Model variety and frontier access | 15% | 5 | 4 | 5 | 5 | 2 |
| Agent harness and repo workflow | 25% | 5 | 4 | 3 | 4 | 5 |
| Steering, verification, and safeguards | 20% | 5 | 4 | 3 | 5 | 5 |
| UX and rollout friction | 10% | 5 | 5 | 3 | 3 | 3 |
| Quota and cost legibility | 10% | 3 | 4 | 3 | 4 | 3 |
| Governance at self-serve team tier | 10% | 4 | 4 | 2 | 3 | 4 |
| Central billing and administration | 10% | 5 | 5 | 2 | 4 | 5 |
| Weighted total | 100% | 4.70 | 4.20 | 3.10 | 4.15 | 4.05 |
| Dimension | Cursor Teams | VS Code + Copilot Business | VS Code + custom model | Kilo Teams | Claude Code Team |
|---|---|---|---|---|---|
| Harness | AI-native VS Code fork; integrated agent, Tab, cloud agents, rules/skills, Bugbot | Mature extension plus GitHub-native agents, reviews, CLI, and broad IDE support | VS Code chat/tools remain available, but semantic search, inline suggestions, and embeddings require Copilot | Open-source extension/CLI with agent modes, cloud agents, and broad provider routing | Terminal-native agent plus IDE integration; deeply optimized around Claude models |
| Models | OpenAI, Anthropic, Google, Cursor/SpaceXAI Grok, Composer, and Auto routing | Broad hosted catalog across OpenAI, Anthropic, and Google; model policy controls | Built-in providers or compatible custom endpoint, including MiniMax-style OpenAI-compatible APIs | 500+ models across 60+ providers, plus BYOK and local models | Claude-only: Fable, Opus, Sonnet, and Haiku; less breadth but tight model–harness co-design |
| Included usage | Per-user Cursor Models and Other Models pools; Premium gives 5× Standard; on-demand overage | 1,900 AI credits per Business seat pooled at billing entity; $0.01/credit overage | No Copilot chat allowance required; usage follows provider price and rate limits | $15/user/month platform fee; inference and cloud compute are separate | Team Standard $20 annual / $25 monthly; Premium $100 / $125 with 5× usage; Code shares the Claude usage pool |
| Heavy-user handling | Mix $40 Standard and $120 Premium seats; upgrade users individually | Shared credit pool lets heavy users consume unused allowance; budgets can cap overage | Use provider limits, gateway budgets, or separate keys | Shared BYOK/balance on Teams; provider-priced inference; Enterprise adds budgets and stronger controls | Mix Standard and Premium seats; paid plans can continue with usage credits at API rates |
| Governance | Teams: SSO, privacy enforcement, analytics, central billing. Enterprise: SCIM, audit, model/MCP/network controls | Organization licensing, policies, budgets, no business-data training; deeper enterprise controls in GitHub Enterprise | Local BYOK is user-managed unless centrally configured; provider terms govern retention | Teams: analytics, billing, privacy controls. Enterprise: SSO/SCIM, audit, allowlists, sandboxing | Team: SSO, central billing, connector controls, no training by default. Enterprise: SCIM, audit, retention, network controls |
| Main trade-off | Requires adopting a dedicated VS Code-derived editor; team usage is not pooled | Best GitHub fit, but agent depth must be proven on your repos and new self-serve Business sign-up is paused for some org plans | Maximum freedom, but fragmented keys, support, quality, and cost ownership | Maximum openness, but more configuration and two-part platform/inference economics | Excellent native Claude harness, but no multi-vendor model choice and a more terminal-centric interaction model |
Correctness depends on instructions being present at the right scope, tools exposing enough evidence, and deterministic controls forcing verification. Scores above summarize this stack; the evidence below shows the mechanics.
| Steering layer | Cursor | Copilot / VS Code | Kilo | Claude Code |
|---|---|---|---|---|
| Base prompt and model adaptation | Cursor says it tunes instructions and tools for every supported frontier model | Strong common agent experience, but BYOK quality depends on how well a custom model follows VS Code's tool schema | Model-agnostic system with custom modes; breadth increases the need to validate each model/tool pairing | Single-vendor co-design gives Anthropic tight control over prompt, tools, compaction, and model behavior |
| Persistent instructions | Enforceable Team Rules, project `.cursor/rules`, AGENTS.md, user rules, and versioned skills/custom modes | Organization, repository, path-specific and AGENTS.md instructions; prompt files, skills, and custom agents | Global config plus shared custom modes/agents with prompts and file/command-scoped permissions | CLAUDE.md at session start, project/user settings, skills, plugins, and persistent subagent memory |
| Repository context | Semantic codebase search, explicit file reads, rules, and model-specific context orchestration | Workspace context plus Copilot semantic search and embeddings; those service features require a Copilot plan | Search/read tools across the worktree; context quality varies with selected model and mode | Agentic search and file reads with compaction; no separate model catalog or router |
| Tools and verification loop | Search, read/edit, terminal, browser, web, MCP, tests, lints, checkpoints, and cloud agents | Read/edit/search/terminal, GitHub issue-to-PR, code review, MCP, hooks, and cloud agent | Read/edit/bash/web/MCP, plan, skills, tasks/subagents, and explicit approval dock | Read/edit/bash/web/MCP, IDE and terminal workflows, tests, hooks, and headless automation |
| Delegation | Built-in Explore, Bash, and Browser subagents isolate noisy context; custom subagents and cloud runs are available | Custom agents and subagents in VS Code/CLI; cloud agents handle GitHub tasks | Task tool launches subagents; modes can be primary or subagent-only | Rich per-subagent controls for tools, model, effort, permissions, MCP, hooks, turns, skills, isolation, and memory |
| Risk gating | Auto-review classifier, deterministic allowlists, shell sandbox, MCP/terminal permissions, and Enterprise policy | Permission prompts and policies; cloud MCP tools can run autonomously, so GitHub recommends read-only allowlists | Every tool defaults to Ask; ordered Allow/Ask/Deny rules, sensitive `.env` protection, and per-agent permissions | Allow/Ask/Deny rules, classifier-backed Auto mode, protected paths, OS-level filesystem/network sandboxing |
| Deterministic guardrails | Command or LLM-evaluated hooks can block/modify prompts, tools, shell, MCP, reads, edits, and completion | Lifecycle hooks can format, scan secrets, audit, or approve/deny tool execution; support varies by surface | Permission rules plus a doom-loop safeguard that pauses repeated non-progress | Hooks run at lifecycle/tool events; unlike prompt guidance, they guarantee checks execute |
| Recovery and human control | Local checkpoints, Git, live steer-at-next-tool-call, cancellation, and review surfaces | Diff review, PR workflow, permission prompts, and Git history | Approve once/always/deny, plan mode, worktree-scoped paths, and repeated-failure pause | Manual/Auto permission modes, sandbox, Git, interruption, max-turn limits, and isolated subagents |
Best fit
Teams willing to standardize on a VS Code-derived editor and use agents for multi-file implementation, debugging, and review.
Validate the undisclosed Standard included-usage amount and real overage on your model mix during the pilot.
Best fit
GitHub-centered organizations that need broad IDE coverage, pooled usage, native licensing, and minimal editor migration.
Budget using the standard 1,900 credits/seat. The temporary 3,000-credit promotion for existing customers ends 1 Sep 2026.
Best fit
Teams with provider contracts, local-model requirements, or a strong preference for open-source and model independence.
Treat model freedom as an architecture choice: it shifts cost, security review, and support responsibility back to your team.
Best fit
Terminal-oriented teams that prioritize a deeply integrated model–harness pair, rich subagent controls, deterministic hooks, and strong sandboxing.
Team seats also include Claude chat and Cowork. This creates product overlap with ChatGPT Business and should be evaluated as a suite alternative, not only a coding add-on.
Rollout and measurement
The complete 10-week pilot, metrics, decision gates, surveys, seat policy, and enterprise triggers.The goal is not to demonstrate that AI can write text or code. The team already knows that. The pilot should prove repeatable workflow gains under company policy, with enough usage data to assign the right seat tier.
Proposed trial mix: ChatGPT 8 Standard + 2 Premium ($450/month on monthly pricing); Cursor 4 Standard + 1 Premium ($280/month). Excludes taxes and usage overage.
| Phase | Weeks | Action | Exit evidence |
|---|---|---|---|
| Baseline | 0 | Each participant nominates two recurring tasks; record current time, quality, and rework | 20 ChatGPT baseline tasks and 10 matched coding-task pairs defined |
| Configure | 1 | Create managed workspaces, enforce privacy, review connectors/models, set budgets | 100% access, policy acknowledgment, connector/model review, and named metric owners |
| Run | 2–6 | Pilot ChatGPT Business broadly; test Cursor, Copilot, and Claude Code on matched engineering tasks; retain Kilo/BYOK as a specialist lane | By week 4: ≥8/10 ChatGPT and ≥4/5 Cursor weekly active; no severe incident |
| Right-size | 4 and 7 | Promote sustained heavy users; downgrade inactive or light users; review overage | Premium retained only for users above 1.5× Standard median usage or with repeated limit pressure |
| Decide | 8–10 | Compare outcomes, governance gaps, support burden, and total cost | Scale only if quality is not worse, safety gates pass, and validated benefit is ≥3× total cost |
| Instrument | Exactly what to capture | Owner and cadence |
|---|---|---|
| 60-second task log | User, task type, tool/model, baseline minutes, actual minutes, outcome, artifact/PR link, manual-fix minutes, and any limit or incident | Participant after every sampled task |
| Matched-task design | Compare the same person's similar recurring tasks before/after; Cursor uses at least two matched pairs per developer | Pilot lead defines at week 0 |
| System evidence | Workspace activity/usage, credits or overage, first CI run, review comments, PR timestamps, and permission/hook events | Admins export every Friday |
| Quality audit | Manager or reviewer blindly checks 20% of sampled deliverables against a task-specific acceptance checklist | Independent reviewer weekly |
| Value calculation | Validated benefit = manager-confirmed hours saved × loaded hourly cost; benefit ratio = validated benefit ÷ seat plus overage cost | Finance or pilot owner at weeks 4 and 10 |
| Metric | Concrete measurement | Week 4 | Week 10 |
|---|---|---|---|
| Weekly active users | Admin analytics plus ≥1 logged task that produced a usable artifact | ≥8 of 10 | ≥8 of 10 in 3 of the final 4 weeks |
| Completed work sessions | Count logs with a deliverable link and reviewer/user acceptance | ≥20 cumulative | ≥60 cumulative, with ≥4 per participant |
| Time saved | Median (baseline minutes − actual minutes) ÷ baseline across matched tasks; manager validates estimates | ≥15% over ≥20 samples | ≥25% over ≥50 samples |
| Company-knowledge quality | Blind review of sampled answers for correct citation and actionability | ≥80% citation-correct over 10 answers | ≥90% citation-correct and ≥75% actionable over 20 answers |
| Limits and overage | Admin export of limit hits, credits, overage, and usage by seat type | Identify users with repeated pressure | Overage ≤20% of license cost; each Premium user >1.5× Standard median usage |
| Policy exceptions | Incident register for prohibited data, unapproved connector/action, or incorrect sharing | 0 severe; ≤2 low-risk | 0 severe, 0 repeated issue, 100% closed |
| Metric | Concrete measurement | Week 4 | Week 10 |
|---|---|---|---|
| Weekly active developers | Cursor analytics plus ≥1 accepted AI-assisted task or PR | ≥4 of 5 | ≥4 of 5 in 3 of the final 4 weeks |
| Lead time | Issue start to review-ready PR for matched task type and developer | ≥15% lower over 10 pairs | ≥25% lower over 20 pairs |
| First-pass CI | Share of AI-assisted PRs whose first full CI run passes | ≥70% and not below baseline | ≥80% and not below baseline |
| Review revisions | Substantive change-request rounds per PR, excluding style-only comments | No worse than baseline | ≥20% fewer than baseline |
| Agent completion rate | Task accepted with all tests passing and <30 minutes of manual repair | ≥60% | ≥70% |
| Cost per accepted task | Allocated seat cost + overage divided by accepted sampled tasks | Establish baseline; ≤$25 directional | ≤$20 and validated benefit ratio ≥3× |
| Security / permission incidents | Review permission, hook, network, secret, and destructive-command events | 0 severe; 100% risky calls reviewed | 0 severe, 0 repeated issue, 100% closed |
| Gate | Scale | Adjust and extend 4 weeks | Stop |
|---|---|---|---|
| Adoption | ChatGPT ≥8/10 and Cursor ≥4/5 sustained | 50–79% active with a fixable enablement gap | <50% active after training and workflow support |
| Productivity | Median matched-task time improves ≥25% | 10–24% improvement or insufficient sample | <10% improvement after ≥20 matched samples |
| Quality | CI, citation correctness, and review rework meet targets | One metric misses by <10 points with a clear intervention | Material regression versus baseline |
| Economics | Validated benefit ÷ total cost ≥3× | 1.5–2.9× with improving trend | <1.5× or uncontrolled overage |
| Safety | 0 severe incidents and all low-risk issues closed | Low-risk isolated issue with completed remediation | Any unresolved severe incident or repeated policy breach |
| Wave | Purpose | What changes after the survey |
|---|---|---|
| Week 2 | Detect onboarding friction and establish perceived quality/steering baselines | Run targeted training, fix access/connectors, and clarify which paid features to try |
| Week 4 | Test whether usage is becoming repeatable and whether seat tiers fit | Reassign Premium seats, adjust rules/prompts, and focus the remaining pilot on proven workflows |
| Week 10 | Measure sustained value, trust, workflow fit, and desire to continue | Combine sentiment with system evidence for scale, extend, or stop decision |
| ID | Exact question | Response | Ask |
|---|---|---|---|
| C1 | In the last 7 days, how many ChatGPT work sessions did you complete? Count one continuous task as one session. | 0 · 1–2 · 3–5 · 6–10 · 11+ | W2, W4, W10 |
| C2 | Which work did you use it for? | Multi-select: writing · research · analysis · company knowledge · files/data · presentations · coding · other | W2, W4, W10 |
| C3 | Overall, how good were the outputs for your real work? | 1 unusable · 2 major rewrite · 3 usable with substantial edits · 4 minor edits · 5 ready to use | W2, W4, W10 |
| C4 | What share of outputs were usable with no more than minor edits? | 0–20% · 21–40% · 41–60% · 61–80% · 81–100% | W2, W4, W10 |
| C5 | Estimate total time saved in the last 7 days. Give one task example with before/after minutes. | 0 · <30 min · 0.5–1 h · 1–2 h · 2–4 h · 4+ h + short example | W2, W4, W10 |
| C6 | Which integrations or company-knowledge sources did you use, and in roughly how many sessions? | Multi-select enabled apps/connectors + 0 · 1–2 · 3–5 · 6+ sessions | W2, W4, W10 |
| C7 | Which paid reasoning options did you deliberately use? | Sol Medium · High · Extra High · Sol Pro · automatic/unsure · none; add approximate session count | W2, W4, W10 |
| C8 | How many Deep Research tasks did you run, and how useful were the final reports? | Count + quality 1–5 + optional report link | W2, W4, W10 |
| C9 | Which other paid capabilities did you use? | Multi-select: Work · Agent mode · shared Projects · scheduled tasks · custom GPTs · Codex · none | W2, W4, W10 |
| C10 | What most limited value this week? | Choose up to 2: inaccurate output · slow · limits · missing source/integration · policy uncertainty · feature unclear · no suitable task · other | W2, W4, W10 |
| C11 | If paid ChatGPT were removed tomorrow, how much would your work be affected, and should the company continue it? | Impact 1–5 · Continue yes/no/unsure · one-sentence reason | W10 only |
| ID | Exact question | Response | Ask |
|---|---|---|---|
| R1 | In the last 7 days, how many Cursor Agent sessions did you run on real repository work? | 0 · 1–2 · 3–5 · 6–10 · 11+ | W2, W4, W10 |
| R2 | Which tasks did you attempt? | Multi-select: bug fix · feature · refactor · tests · debugging · code review · docs · investigation · other | W2, W4, W10 |
| R3 | Which models or routing modes did you use most? | Auto · Grok · Composer · Claude · GPT · Gemini · Kimi · other/unsure | W2, W4, W10 |
| R4 | How good was the final code or technical output? | 1 unusable · 2 major rewrite · 3 substantial edits · 4 minor edits · 5 accepted as produced | W2, W4, W10 |
| R5 | How much manual steering did the agent need to reach the correct result? | 1 none · 2 one correction · 3 several corrections · 4 frequent steering · 5 constant supervision | W2, W4, W10 |
| R6 | For a typical accepted task, how much manual repair was needed after the agent stopped? | 0 · <15 min · 15–30 min · 31–60 min · >60 min | W2, W4, W10 |
| R7 | What share of attempted tasks reached the requested result with tests passing? | 0–20% · 21–40% · 41–60% · 61–80% · 81–100% | W2, W4, W10 |
| R8 | Estimate total time saved in the last 7 days. Give one task example with before/after minutes. | 0 · <30 min · 0.5–1 h · 1–2 h · 2–4 h · 4+ h + PR/task example | W2, W4, W10 |
| R9 | Which Cursor capabilities materially helped? | Multi-select: Tab · Agent · terminal · browser · MCP · rules · skills · subagents · cloud agents · Bugbot · checkpoints | W2, W4, W10 |
| R10 | How often did you hit a failed loop, revert a change, or abandon the agent and finish manually? | 0 · 1 · 2–3 · 4–5 · 6+; optional example | W2, W4, W10 |
| R11 | Did permissions or safeguards feel appropriate? | Too restrictive · about right · too permissive · unsure; describe any unsafe or blocked action | W2, W4, W10 |
| R12 | If Cursor were removed tomorrow, how much would your work be affected, and should the company continue it? | Impact 1–5 · Continue yes/no/unsure · one-sentence reason | W10 only |
| Survey signal | Week 2 | Week 4 | Week 10 |
|---|---|---|---|
| Response rate | 100%: 10 ChatGPT + 5 Cursor | 100% | 100% |
| ChatGPT sessions | ≥7/10 report 2+ sessions | ≥8/10 report 3+ sessions | ≥8/10 report 3+ sessions and sustained system activity |
| ChatGPT perceived quality | Median ≥3.2/5 | Median ≥3.5/5 | Median ≥4.0/5; ≥70% outputs need only minor edits |
| ChatGPT time and paid-feature use | Median ≥0.5 h saved; ≥6/10 tried Sol or one paid feature | Median ≥1 h; ≥8/10 use Sol and ≥4/10 use Deep Research or an integration | Median ≥2 h; ≥8/10 use a paid feature weekly and ≥6/10 use integrations |
| Cursor sessions | ≥4/5 report 2+ sessions | ≥4/5 report 3+ sessions | ≥4/5 report 3+ sessions and sustained system activity |
| Cursor quality and steering | Quality median ≥3.2; steering median ≤3.5 | Quality ≥3.5; steering ≤3.0 | Quality ≥4.0; steering ≤2.5; ≥70% tasks complete with <30 min repair |
| Cursor time and advanced-feature use | Median ≥0.5 h saved; all 5 use Agent + terminal | Median ≥1.5 h; ≥3/5 use rules, skills, browser, MCP, or subagents | Median ≥2 h; ≥4/5 rely on at least 2 advanced capabilities |
| Continuation signal | Not used as a gate | Ask informally for emerging blockers | ≥8/10 ChatGPT and ≥4/5 Cursor participants answer Continue |
| Tier | Starting rule | Upgrade trigger | Downgrade trigger |
|---|---|---|---|
| Standard | ChatGPT: 8 users. Cursor: 4 developers | Two limit-pressure weeks plus accepted work and >1.5× Standard median usage | Inactive in 2 of 4 weeks or value metrics remain below target |
| Premium | ChatGPT: 2 users. Cursor: 1 developer | Already Premium; add overage budget only when unit economics remain positive | Usage ≤1.5× Standard median or no measurable output advantage at week 7 |
| Specialist BYOK | Only for approved model/locality requirements | A documented capability gap in managed hosted models | Support, privacy, or cost ownership becomes fragmented |
| Need | ChatGPT | Cursor | Coding alternatives |
|---|---|---|---|
| SCIM / automated lifecycle | Enterprise | Enterprise | Claude/Kilo Enterprise or GitHub enterprise controls |
| Fine-grained model / connector / MCP policy | Enterprise for deeper RBAC | Enterprise | Copilot org/enterprise policy or Kilo Enterprise allowlists |
| Audit logs / service accounts | Enterprise | Enterprise | Copilot/GitHub audit, Claude Enterprise, or Kilo Enterprise |
| Invoice / PO / negotiated terms | Contracted offering | Enterprise | Vendor enterprise plan |
| Pooled coding usage | Not the coding comparison | Enterprise | Copilot Business already pools AI credits |