Complete analytical reference

From free AI chat
to managed AI work

The full business case behind the presentation: economics, ChatGPT plan differentiation, current model benchmarks, coding-harness mechanics, governance, and a measurable 10-week rollout.

Business case v2 Verified 26 Aug 2026 Converted from the original Cursor Canvas
Readable reference generated from the complete source canvas This page preserves every section, table, survey question, and cited source from the original 1,867-line analytical artifact.
Interactive seat calculator
01

Decision and economics

Recommendation, interactive-seat assumptions, alternatives, and decision boundaries.

The business case in one page

The step beyond free chat is not only “better answers.” It is a managed operating environment: dependable capacity, frontier-model access, company context, enforceable data policy, usage visibility, and one bill.

LayerBuyWhy nowWho gets it
General AIChatGPT BusinessCompany knowledge, higher limits, secure workspace, admin controlsKnowledge workers who already use AI chat
Software deliveryCursor TeamsAI-native agent harness, broad frontier models, repo context, team policyDevelopers and technical builders
Illustrative seat mixAdjustable

ChatGPT Business

Paid users

10

Premium / heavy users

2

Cursor Teams

Paid users

5

Premium / heavy users

1

ChatGPT billing

AnnualMonthly

$450ChatGPT / month
$280Cursor / month
$730Combined / month

ChatGPT: 8 Standard + 2 Premium. Cursor: 4 Standard + 1 Premium. Estimate excludes tax, usage overages, and API purchases.

Why mixed seat levels matter

ChatGPT Business5× usage

Standard: $25/user/month. Premium: $125/user/month with 5× usage and no five-hour usage limit.

Standard and Premium seats can coexist in one workspace and be assigned or reassigned. This directly supports a “baseline + power-user” policy.

Official seat documentation
Cursor Teams5× usage

Standard: $40/user/month. Premium: $120/user/month with 5× the Standard Agent allowance.

Admins can upgrade or downgrade individual users. Upgrades are immediate and prorated; included usage remains per-user rather than pooled on Teams.

Official team pricing

Decision boundaries

If this is the priorityBest starting pointReason
Fastest path beyond free chatChatGPT BusinessSmall migration burden; immediate gains in capacity, company context, and governance
Best integrated coding-agent experienceCursor TeamsAI-native editor, multi-file agent loop, model routing, shared rules and skills
Lowest-friction GitHub / multi-IDE rolloutCopilot BusinessNative GitHub administration and broad IDE reach
Strongest terminal-native steering and isolationClaude Code TeamRich subagents, hooks, permissions, sandboxing, and a tightly co-designed Claude harness
Maximum model freedom or local/BYOKKilo or VS Code BYOKHundreds of providers/models, local endpoints, and direct provider billing
02

ChatGPT business case

Free-versus-Business comparison, GPT-5.6 model evidence, paid capabilities, and governance.

ChatGPT: free chat versus a managed AI workspace

The strongest argument is operational, not cosmetic. Business turns fragmented personal use into an environment the company can fund, govern, connect, and measure.

DimensionFree personal workspaceChatGPT BusinessCase to make
Usage and continuityLimited messages, uploads, deep research, memory, context, Codex, and slower image generationHigher included limits; Standard or 5× Premium seats; workspace credits can extend usageFewer interruptions during real work
Model access and qualityGPT-5.6 Luna is the current default; no access to Sol reasoning tiersGPT-5.6 Sol at Medium, High and Extra High, plus GPT-5.6 Sol Pro; Luna/Terra/Sol in Work and CodexHarder work can use stronger reasoning, not only a faster chatbot
Company knowledgePersonal connections and ad hoc file uploads; no shared company workspace policyCompany knowledge across enabled apps such as Slack, SharePoint, Drive, GitHub, HubSpot, and Asana, with citationsAnswers can be grounded in current internal sources
Connector governanceConfigured by each individualAdmins manage plugins/apps and permissions; access respects each user's source-system permissionsConnectors become governable rather than shadow integrations
Data handlingPersonal-workspace training sharing is enabled by default, with user opt-outWorkspace inputs and outputs are excluded from model training by defaultA safer default for company information
AdministrationNo organization-level roles, usage visibility, or policyOwners/admins, centralized billing, usage analytics, spend controls, MFA and SAML SSOOne accountable owner and policy surface
Right-sizingOne free tier per personMix and reassign Standard ($20 annual / $25 monthly) and Premium ($100 annual / $125 monthly)Pay for intensity where it is proven

What the GPT names actually mean in ChatGPT

ModelPositioningWhere users encounter itQuality evidence
GPT-5.5 InstantFast conversational baseline and previous defaultReplaced GPT-5.3 Instant for all users in May 2026; now being superseded by the 5.6 rollout52.5% fewer hallucinated claims than GPT-5.3 Instant on OpenAI's high-stakes internal prompts; no direct same-harness 5.6 comparison published
GPT-5.6 LunaFastest, lowest-cost 5.6 modelDefault and Think for Free/Go; Work and Codex on eligible paid plansGood high-volume default; lower peak reasoning than Terra and Sol
GPT-5.6 TerraBalanced capability, speed, and costWork and Codex; not selectable in ordinary ChatGPT conversationsNear-Sol coding performance in OpenAI's published model-level evaluations
GPT-5.6 SolFlagship for complex reasoning, research, coding, and long-running workInstant with automatic reasoning plus Medium, High and Extra High on BusinessBest 5.6 family result across the three engineering benchmarks below
GPT-5.6 Sol ProHighest-capability ChatGPT option for difficult, longer-running workflowsPro picker option on Business, Enterprise, and individual ProNo clean public ChatGPT-specific apples-to-apples score located; do not substitute the API's Sol Ultra result

OpenAI-published model quality snapshot

Higher is better. GPT-5.5 here means the reasoning/Thinking model, not GPT-5.5 Instant. These are model-level evaluations reported by OpenAI, not a guarantee of ChatGPT product results.

EvaluationGPT-5.5 ThinkingGPT-5.6 LunaGPT-5.6 TerraGPT-5.6 Sol
AA Intelligence Index v4.1 · index54.851.255.058.9
AA Coding Agent Index v1.1 · index76.474.677.480.0
SWE-Bench Pro · resolved59.4%62.7%63.4%64.6%
DeepSWE v1.1 · resolved67.0%67.2%69.6%72.7%
Terminal-Bench 2.1 · pass85.6%84.7%87.4%88.8%
BrowseComp · solved84.4%83.3%87.5%90.4%
OSWorld 2.0 · success47.5%45.6%50.2%62.6%

Source: OpenAI GPT-5.6 launch report · June 2026. Results depend on reasoning effort and harness configuration.

Capabilities beyond free chat

CapabilityFreeChatGPT BusinessProcurement meaning
Deep researchAvailable with a small plan-dependent allowanceExpanded use; editable research plan, files, web, selected sites, and enabled apps with cited reportsA real research workflow, but capacity—not mere availability—is the paid distinction
Agent modeUnavailable40 initiated messages per month before flexible overage, subject to current limitsDelegates browser-and-tool workflows rather than only producing an answer
ProjectsAvailable; 5 files per projectShared team projects, up to 40 files per project, and workspace privacy defaultsBusiness adds collaboration and governed data handling
Scheduled tasksLimited cadence and flexible delivery windows; up to 3 activeHigher active-task limit, hourly/exact schedules, and eligible event triggers in Work; current help text conflicts between 10 and 15 active tasksUseful for monitoring and recurring work, not just reminders
Create custom GPTsCannot create or publish new GPTs on personal plansMembers can create managed workspace GPTs when admins permitReusable internal assistants become company-managed assets
Company knowledgeNo managed organization knowledge layerPermission-aware answers across approved apps, with source citations; currently web-onlyThe clearest capability that is genuinely Business-specific
Workspace agentsNot available as a managed workspace capabilityCreate, preview, share, and schedule repeatable agents within the workspaceMoves from one-off prompts toward repeatable workflows
ChatGPT WorkLimited accessMulti-step agent for research, connected apps/files, and finished documents, sheets, presentations, reports, and SitesDelegated deliverables rather than chat responses
CodexLimited access; Terra where availableExpanded coding access with Sol, Terra, and Luna plus workspace controlsUseful secondary coding surface, though Cursor remains the dedicated editor proposal

Frame connectors accurately

Call this “permission-aware company knowledge,” not a universal data warehouse. Quality depends on which apps are enabled, source freshness, indexing mode, and the user's existing access.

Know when Business is not enough

Move the conversation to Enterprise when the organization needs SCIM, EKM, domain verification, fine-grained RBAC, custom retention, invoice/PO terms, or more than the self-serve Business workspace limit.

Suggested executive wording

30-second case

“The team already uses free AI, so the choice is not adoption versus no adoption. It is unmanaged personal use versus a secure company workspace. ChatGPT Business gives us stronger models and dependable capacity, grounds answers in approved company sources, excludes workspace data from training by default, and consolidates billing and administration. Mixed Standard and Premium seats let us scale heavy users without overbuying for everyone.”

Official evidence for ChatGPT claims
03

Model benchmarks

CursorBench, Artificial Analysis, OpenRouter caveats, cost, and model–harness interpretation.

Model quality: use more than one benchmark

Coding quality is a property of the model–harness pair. CursorBench is useful for comparing models inside Cursor's own agent loop, while Artificial Analysis compares complete coding-agent systems. Neither replaces a matched pilot on your repositories.

CursorBench 3.2 · best published configuration per model

Higher is better. CursorBench uses ambiguous, multi-file tasks derived from real Cursor sessions. It is first-party and not independently reproducible; effort levels and cost per task vary materially. Cursor warns that small gaps may not be statistically meaningful, so the 0.8-point spread across the top three is not decision-grade.

RankModelBest scoreInterpretation
1Grok 4.6 · Extra High70.8%Current peak; $2.81/task in Cursor's published run
2Claude Fable 5 · Max70.5%Near peak, but $17.32/task
3Claude Opus 5 · Max70.0%Peak Claude value tier; $8.23/task
4GPT-5.6 Sol · Max67.2%Strong flagship result; $5.69/task
5GPT-5.6 Terra · Max64.9%Balanced 5.6 option
6Gemini 3.7 Flash61.6%High-throughput frontier option
7Claude Sonnet 561.5%Lower-cost Claude default
8GPT-5.6 Luna61.1%Fastest 5.6 tier
9Kimi K360.8%Open-weight, long-horizon option
10GPT-5.558.4%Previous OpenAI reasoning generation
11Composer 2.556.1%Lower peak score but $0.44/task and Cursor-native tuning

Source: CursorBench 3.2 public leaderboard · snapshot checked 26 Aug 2026. The Grok 4.5 training-data caveat is omitted because this table uses Grok 4.6.

Read score and cost together

The top 3.6 percentage points span roughly $2.81 to $17.32 per task. Composer 2.5 scores 14.7 points below the leader but costs about one-sixth as much as Grok 4.6 Extra High and about one-fortieth as much as Fable 5 Max in Cursor's run.

Do not infer product quality from one row

CursorBench holds Cursor's harness broadly constant to expose model differences. It does not tell us how Cursor versus Claude Code versus Copilot performs with each product's native prompting, context management, safeguards, and verification loops.

Independent model–harness cross-check

Artificial Analysis Coding Agent Index v1.4: simple average of DeepSWE, Terminal-Bench 2.1, and SWE-Atlas-QnA pass@1. These rows compare complete systems, not isolated models.

Agent harnessRepresentative modelIndexCost / taskWhat it shows
Claude CodeClaude Opus 5 · xhigh68$8.17Highest of this current cross-check; strongest Terminal-Bench and repo-Q&A mix
CodexGPT-5.6 Sol · max65$6.42Best DeepSWE result here and materially faster than the Claude Code run
Grok BuildGrok 4.5 · high64$2.44Near-frontier composite at much lower cost
Kimi Code CLIKimi K363$3.08Strong open-weight model–native harness pairing
OpenCodeGemini 3.7 Flash · high60$1.27Strong terminal score and the lowest task cost in this set
Cursor CLIGPT-5.5 · medium47$2.00An older Cursor/model configuration; absence of current Cursor models prevents a fair product ranking
Benchmark sources and methodology
04

Coding harness assessment

Cursor, Copilot, VS Code BYOK, Kilo, and Claude Code across quality, steering, safety, and cost.

Coding harness: evaluate the system, not just the model

Cursor, Copilot, VS Code BYOK, Kilo, and Claude Code increasingly overlap on frontier-model capability. The durable differentiator is how the harness builds the base prompt, gathers repository context, exposes tools, lets teams steer behavior, verifies results, recovers from failure, and gates risky actions.

Weighted decision score

Internal assessment, 1–5. It is a hypothesis for the pilot, not a vendor benchmark. Weights favor end-to-end agent work over autocomplete.

CriterionWeightCursor TeamsCopilot BusinessVS Code + BYOKKilo TeamsClaude Code Team
Model variety and frontier access15%54552
Agent harness and repo workflow25%54345
Steering, verification, and safeguards20%54355
UX and rollout friction10%55333
Quota and cost legibility10%34343
Governance at self-serve team tier10%44234
Central billing and administration10%55245
Weighted total100%4.704.203.104.154.05

What sits behind the score

DimensionCursor TeamsVS Code + Copilot BusinessVS Code + custom modelKilo TeamsClaude Code Team
HarnessAI-native VS Code fork; integrated agent, Tab, cloud agents, rules/skills, BugbotMature extension plus GitHub-native agents, reviews, CLI, and broad IDE supportVS Code chat/tools remain available, but semantic search, inline suggestions, and embeddings require CopilotOpen-source extension/CLI with agent modes, cloud agents, and broad provider routingTerminal-native agent plus IDE integration; deeply optimized around Claude models
ModelsOpenAI, Anthropic, Google, Cursor/SpaceXAI Grok, Composer, and Auto routingBroad hosted catalog across OpenAI, Anthropic, and Google; model policy controlsBuilt-in providers or compatible custom endpoint, including MiniMax-style OpenAI-compatible APIs500+ models across 60+ providers, plus BYOK and local modelsClaude-only: Fable, Opus, Sonnet, and Haiku; less breadth but tight model–harness co-design
Included usagePer-user Cursor Models and Other Models pools; Premium gives 5× Standard; on-demand overage1,900 AI credits per Business seat pooled at billing entity; $0.01/credit overageNo Copilot chat allowance required; usage follows provider price and rate limits$15/user/month platform fee; inference and cloud compute are separateTeam Standard $20 annual / $25 monthly; Premium $100 / $125 with 5× usage; Code shares the Claude usage pool
Heavy-user handlingMix $40 Standard and $120 Premium seats; upgrade users individuallyShared credit pool lets heavy users consume unused allowance; budgets can cap overageUse provider limits, gateway budgets, or separate keysShared BYOK/balance on Teams; provider-priced inference; Enterprise adds budgets and stronger controlsMix Standard and Premium seats; paid plans can continue with usage credits at API rates
GovernanceTeams: SSO, privacy enforcement, analytics, central billing. Enterprise: SCIM, audit, model/MCP/network controlsOrganization licensing, policies, budgets, no business-data training; deeper enterprise controls in GitHub EnterpriseLocal BYOK is user-managed unless centrally configured; provider terms govern retentionTeams: analytics, billing, privacy controls. Enterprise: SSO/SCIM, audit, allowlists, sandboxingTeam: SSO, central billing, connector controls, no training by default. Enterprise: SCIM, audit, retention, network controls
Main trade-offRequires adopting a dedicated VS Code-derived editor; team usage is not pooledBest GitHub fit, but agent depth must be proven on your repos and new self-serve Business sign-up is paused for some org plansMaximum freedom, but fragmented keys, support, quality, and cost ownershipMaximum openness, but more configuration and two-part platform/inference economicsExcellent native Claude harness, but no multi-vendor model choice and a more terminal-centric interaction model

How each harness steers toward a correct result

Correctness depends on instructions being present at the right scope, tools exposing enough evidence, and deterministic controls forcing verification. Scores above summarize this stack; the evidence below shows the mechanics.

Steering layerCursorCopilot / VS CodeKiloClaude Code
Base prompt and model adaptationCursor says it tunes instructions and tools for every supported frontier modelStrong common agent experience, but BYOK quality depends on how well a custom model follows VS Code's tool schemaModel-agnostic system with custom modes; breadth increases the need to validate each model/tool pairingSingle-vendor co-design gives Anthropic tight control over prompt, tools, compaction, and model behavior
Persistent instructionsEnforceable Team Rules, project `.cursor/rules`, AGENTS.md, user rules, and versioned skills/custom modesOrganization, repository, path-specific and AGENTS.md instructions; prompt files, skills, and custom agentsGlobal config plus shared custom modes/agents with prompts and file/command-scoped permissionsCLAUDE.md at session start, project/user settings, skills, plugins, and persistent subagent memory
Repository contextSemantic codebase search, explicit file reads, rules, and model-specific context orchestrationWorkspace context plus Copilot semantic search and embeddings; those service features require a Copilot planSearch/read tools across the worktree; context quality varies with selected model and modeAgentic search and file reads with compaction; no separate model catalog or router
Tools and verification loopSearch, read/edit, terminal, browser, web, MCP, tests, lints, checkpoints, and cloud agentsRead/edit/search/terminal, GitHub issue-to-PR, code review, MCP, hooks, and cloud agentRead/edit/bash/web/MCP, plan, skills, tasks/subagents, and explicit approval dockRead/edit/bash/web/MCP, IDE and terminal workflows, tests, hooks, and headless automation
DelegationBuilt-in Explore, Bash, and Browser subagents isolate noisy context; custom subagents and cloud runs are availableCustom agents and subagents in VS Code/CLI; cloud agents handle GitHub tasksTask tool launches subagents; modes can be primary or subagent-onlyRich per-subagent controls for tools, model, effort, permissions, MCP, hooks, turns, skills, isolation, and memory
Risk gatingAuto-review classifier, deterministic allowlists, shell sandbox, MCP/terminal permissions, and Enterprise policyPermission prompts and policies; cloud MCP tools can run autonomously, so GitHub recommends read-only allowlistsEvery tool defaults to Ask; ordered Allow/Ask/Deny rules, sensitive `.env` protection, and per-agent permissionsAllow/Ask/Deny rules, classifier-backed Auto mode, protected paths, OS-level filesystem/network sandboxing
Deterministic guardrailsCommand or LLM-evaluated hooks can block/modify prompts, tools, shell, MCP, reads, edits, and completionLifecycle hooks can format, scan secrets, audit, or approve/deny tool execution; support varies by surfacePermission rules plus a doom-loop safeguard that pauses repeated non-progressHooks run at lifecycle/tool events; unlike prompt guidance, they guarantee checks execute
Recovery and human controlLocal checkpoints, Git, live steer-at-next-tool-call, cancellation, and review surfacesDiff review, PR workflow, permission prompts, and Git historyApprove once/always/deny, plan mode, worktree-scoped paths, and repeated-failure pauseManual/Auto permission modes, sandbox, Git, interruption, max-turn limits, and isolated subagents
Cursor Teams

Best fit

Teams willing to standardize on a VS Code-derived editor and use agents for multi-file implementation, debugging, and review.

Validate the undisclosed Standard included-usage amount and real overage on your model mix during the pilot.

Copilot Business

Best fit

GitHub-centered organizations that need broad IDE coverage, pooled usage, native licensing, and minimal editor migration.

Budget using the standard 1,900 credits/seat. The temporary 3,000-credit promotion for existing customers ends 1 Sep 2026.

Kilo / BYOK

Best fit

Teams with provider contracts, local-model requirements, or a strong preference for open-source and model independence.

Treat model freedom as an architecture choice: it shifts cost, security review, and support responsibility back to your team.

Claude Code Team

Best fit

Terminal-oriented teams that prioritize a deeply integrated model–harness pair, rich subagent controls, deterministic hooks, and strong sandboxing.

Team seats also include Claude chat and Cowork. This creates product overlap with ChatGPT Business and should be evaluated as a suite alternative, not only a coding add-on.

Official evidence for coding-tool claims
05

Rollout and measurement

The complete 10-week pilot, metrics, decision gates, surveys, seat policy, and enterprise triggers.

10-week pilot: prove value and right-size seats

The goal is not to demonstrate that AI can write text or code. The team already knows that. The pilot should prove repeatable workflow gains under company policy, with enough usage data to assign the right seat tier.

10ChatGPT participants
5Cursor participants
$730Base licenses / month

Proposed trial mix: ChatGPT 8 Standard + 2 Premium ($450/month on monthly pricing); Cursor 4 Standard + 1 Premium ($280/month). Excludes taxes and usage overage.

PhaseWeeksActionExit evidence
Baseline0Each participant nominates two recurring tasks; record current time, quality, and rework20 ChatGPT baseline tasks and 10 matched coding-task pairs defined
Configure1Create managed workspaces, enforce privacy, review connectors/models, set budgets100% access, policy acknowledgment, connector/model review, and named metric owners
Run2–6Pilot ChatGPT Business broadly; test Cursor, Copilot, and Claude Code on matched engineering tasks; retain Kilo/BYOK as a specialist laneBy week 4: ≥8/10 ChatGPT and ≥4/5 Cursor weekly active; no severe incident
Right-size4 and 7Promote sustained heavy users; downgrade inactive or light users; review overagePremium retained only for users above 1.5× Standard median usage or with repeated limit pressure
Decide8–10Compare outcomes, governance gaps, support burden, and total costScale only if quality is not worse, safety gates pass, and validated benefit is ≥3× total cost

One measurement protocol for both products

InstrumentExactly what to captureOwner and cadence
60-second task logUser, task type, tool/model, baseline minutes, actual minutes, outcome, artifact/PR link, manual-fix minutes, and any limit or incidentParticipant after every sampled task
Matched-task designCompare the same person's similar recurring tasks before/after; Cursor uses at least two matched pairs per developerPilot lead defines at week 0
System evidenceWorkspace activity/usage, credits or overage, first CI run, review comments, PR timestamps, and permission/hook eventsAdmins export every Friday
Quality auditManager or reviewer blindly checks 20% of sampled deliverables against a task-specific acceptance checklistIndependent reviewer weekly
Value calculationValidated benefit = manager-confirmed hours saved × loaded hourly cost; benefit ratio = validated benefit ÷ seat plus overage costFinance or pilot owner at weeks 4 and 10

ChatGPT measurement · 10 people

MetricConcrete measurementWeek 4Week 10
Weekly active usersAdmin analytics plus ≥1 logged task that produced a usable artifact≥8 of 10≥8 of 10 in 3 of the final 4 weeks
Completed work sessionsCount logs with a deliverable link and reviewer/user acceptance≥20 cumulative≥60 cumulative, with ≥4 per participant
Time savedMedian (baseline minutes − actual minutes) ÷ baseline across matched tasks; manager validates estimates≥15% over ≥20 samples≥25% over ≥50 samples
Company-knowledge qualityBlind review of sampled answers for correct citation and actionability≥80% citation-correct over 10 answers≥90% citation-correct and ≥75% actionable over 20 answers
Limits and overageAdmin export of limit hits, credits, overage, and usage by seat typeIdentify users with repeated pressureOverage ≤20% of license cost; each Premium user >1.5× Standard median usage
Policy exceptionsIncident register for prohibited data, unapproved connector/action, or incorrect sharing0 severe; ≤2 low-risk0 severe, 0 repeated issue, 100% closed

Cursor measurement · 5 people

MetricConcrete measurementWeek 4Week 10
Weekly active developersCursor analytics plus ≥1 accepted AI-assisted task or PR≥4 of 5≥4 of 5 in 3 of the final 4 weeks
Lead timeIssue start to review-ready PR for matched task type and developer≥15% lower over 10 pairs≥25% lower over 20 pairs
First-pass CIShare of AI-assisted PRs whose first full CI run passes≥70% and not below baseline≥80% and not below baseline
Review revisionsSubstantive change-request rounds per PR, excluding style-only commentsNo worse than baseline≥20% fewer than baseline
Agent completion rateTask accepted with all tests passing and <30 minutes of manual repair≥60%≥70%
Cost per accepted taskAllocated seat cost + overage divided by accepted sampled tasksEstablish baseline; ≤$25 directional≤$20 and validated benefit ratio ≥3×
Security / permission incidentsReview permission, hook, network, secret, and destructive-command events0 severe; 100% risky calls reviewed0 severe, 0 repeated issue, 100% closed

Decision gates

GateScaleAdjust and extend 4 weeksStop
AdoptionChatGPT ≥8/10 and Cursor ≥4/5 sustained50–79% active with a fixable enablement gap<50% active after training and workflow support
ProductivityMedian matched-task time improves ≥25%10–24% improvement or insufficient sample<10% improvement after ≥20 matched samples
QualityCI, citation correctness, and review rework meet targetsOne metric misses by <10 points with a clear interventionMaterial regression versus baseline
EconomicsValidated benefit ÷ total cost ≥3×1.5–2.9× with improving trend<1.5× or uncontrolled overage
Safety0 severe incidents and all low-risk issues closedLow-risk isolated issue with completed remediationAny unresolved severe incident or repeated policy breach

Pulse surveys at weeks 2, 4, and 10

WavePurposeWhat changes after the survey
Week 2Detect onboarding friction and establish perceived quality/steering baselinesRun targeted training, fix access/connectors, and clarify which paid features to try
Week 4Test whether usage is becoming repeatable and whether seat tiers fitReassign Premium seats, adjust rules/prompts, and focus the remaining pilot on proven workflows
Week 10Measure sustained value, trust, workflow fit, and desire to continueCombine sentiment with system evidence for scale, extend, or stop decision
ChatGPT survey draft11
IDExact questionResponseAsk
C1In the last 7 days, how many ChatGPT work sessions did you complete? Count one continuous task as one session.0 · 1–2 · 3–5 · 6–10 · 11+W2, W4, W10
C2Which work did you use it for?Multi-select: writing · research · analysis · company knowledge · files/data · presentations · coding · otherW2, W4, W10
C3Overall, how good were the outputs for your real work?1 unusable · 2 major rewrite · 3 usable with substantial edits · 4 minor edits · 5 ready to useW2, W4, W10
C4What share of outputs were usable with no more than minor edits?0–20% · 21–40% · 41–60% · 61–80% · 81–100%W2, W4, W10
C5Estimate total time saved in the last 7 days. Give one task example with before/after minutes.0 · <30 min · 0.5–1 h · 1–2 h · 2–4 h · 4+ h + short exampleW2, W4, W10
C6Which integrations or company-knowledge sources did you use, and in roughly how many sessions?Multi-select enabled apps/connectors + 0 · 1–2 · 3–5 · 6+ sessionsW2, W4, W10
C7Which paid reasoning options did you deliberately use?Sol Medium · High · Extra High · Sol Pro · automatic/unsure · none; add approximate session countW2, W4, W10
C8How many Deep Research tasks did you run, and how useful were the final reports?Count + quality 1–5 + optional report linkW2, W4, W10
C9Which other paid capabilities did you use?Multi-select: Work · Agent mode · shared Projects · scheduled tasks · custom GPTs · Codex · noneW2, W4, W10
C10What most limited value this week?Choose up to 2: inaccurate output · slow · limits · missing source/integration · policy uncertainty · feature unclear · no suitable task · otherW2, W4, W10
C11If paid ChatGPT were removed tomorrow, how much would your work be affected, and should the company continue it?Impact 1–5 · Continue yes/no/unsure · one-sentence reasonW10 only
Cursor survey draft12
IDExact questionResponseAsk
R1In the last 7 days, how many Cursor Agent sessions did you run on real repository work?0 · 1–2 · 3–5 · 6–10 · 11+W2, W4, W10
R2Which tasks did you attempt?Multi-select: bug fix · feature · refactor · tests · debugging · code review · docs · investigation · otherW2, W4, W10
R3Which models or routing modes did you use most?Auto · Grok · Composer · Claude · GPT · Gemini · Kimi · other/unsureW2, W4, W10
R4How good was the final code or technical output?1 unusable · 2 major rewrite · 3 substantial edits · 4 minor edits · 5 accepted as producedW2, W4, W10
R5How much manual steering did the agent need to reach the correct result?1 none · 2 one correction · 3 several corrections · 4 frequent steering · 5 constant supervisionW2, W4, W10
R6For a typical accepted task, how much manual repair was needed after the agent stopped?0 · <15 min · 15–30 min · 31–60 min · >60 minW2, W4, W10
R7What share of attempted tasks reached the requested result with tests passing?0–20% · 21–40% · 41–60% · 61–80% · 81–100%W2, W4, W10
R8Estimate total time saved in the last 7 days. Give one task example with before/after minutes.0 · <30 min · 0.5–1 h · 1–2 h · 2–4 h · 4+ h + PR/task exampleW2, W4, W10
R9Which Cursor capabilities materially helped?Multi-select: Tab · Agent · terminal · browser · MCP · rules · skills · subagents · cloud agents · Bugbot · checkpointsW2, W4, W10
R10How often did you hit a failed loop, revert a change, or abandon the agent and finish manually?0 · 1 · 2–3 · 4–5 · 6+; optional exampleW2, W4, W10
R11Did permissions or safeguards feel appropriate?Too restrictive · about right · too permissive · unsure; describe any unsafe or blocked actionW2, W4, W10
R12If Cursor were removed tomorrow, how much would your work be affected, and should the company continue it?Impact 1–5 · Continue yes/no/unsure · one-sentence reasonW10 only

Survey milestone targets

Survey signalWeek 2Week 4Week 10
Response rate100%: 10 ChatGPT + 5 Cursor100%100%
ChatGPT sessions≥7/10 report 2+ sessions≥8/10 report 3+ sessions≥8/10 report 3+ sessions and sustained system activity
ChatGPT perceived qualityMedian ≥3.2/5Median ≥3.5/5Median ≥4.0/5; ≥70% outputs need only minor edits
ChatGPT time and paid-feature useMedian ≥0.5 h saved; ≥6/10 tried Sol or one paid featureMedian ≥1 h; ≥8/10 use Sol and ≥4/10 use Deep Research or an integrationMedian ≥2 h; ≥8/10 use a paid feature weekly and ≥6/10 use integrations
Cursor sessions≥4/5 report 2+ sessions≥4/5 report 3+ sessions≥4/5 report 3+ sessions and sustained system activity
Cursor quality and steeringQuality median ≥3.2; steering median ≤3.5Quality ≥3.5; steering ≤3.0Quality ≥4.0; steering ≤2.5; ≥70% tasks complete with <30 min repair
Cursor time and advanced-feature useMedian ≥0.5 h saved; all 5 use Agent + terminalMedian ≥1.5 h; ≥3/5 use rules, skills, browser, MCP, or subagentsMedian ≥2 h; ≥4/5 rely on at least 2 advanced capabilities
Continuation signalNot used as a gateAsk informally for emerging blockers≥8/10 ChatGPT and ≥4/5 Cursor participants answer Continue

Seat assignment policy

TierStarting ruleUpgrade triggerDowngrade trigger
StandardChatGPT: 8 users. Cursor: 4 developersTwo limit-pressure weeks plus accepted work and >1.5× Standard median usageInactive in 2 of 4 weeks or value metrics remain below target
PremiumChatGPT: 2 users. Cursor: 1 developerAlready Premium; add overage budget only when unit economics remain positiveUsage ≤1.5× Standard median or no measurable output advantage at week 7
Specialist BYOKOnly for approved model/locality requirementsA documented capability gap in managed hosted modelsSupport, privacy, or cost ownership becomes fragmented

Enterprise escalation triggers

NeedChatGPTCursorCoding alternatives
SCIM / automated lifecycleEnterpriseEnterpriseClaude/Kilo Enterprise or GitHub enterprise controls
Fine-grained model / connector / MCP policyEnterprise for deeper RBACEnterpriseCopilot org/enterprise policy or Kilo Enterprise allowlists
Audit logs / service accountsEnterpriseEnterpriseCopilot/GitHub audit, Claude Enterprise, or Kilo Enterprise
Invoice / PO / negotiated termsContracted offeringEnterpriseVendor enterprise plan
Pooled coding usageNot the coding comparisonEnterpriseCopilot Business already pools AI credits