Team AI Proposal
Full report 26 Aug 2026

A practical proposal for the next stage of AI adoption

From free AI chat
to managed AI work

Introduce ChatGPT Business for company-wide knowledge work and Cursor Teams for software delivery—with mixed seat levels, measurable outcomes, and clear guardrails.

Use to navigate
10 ChatGPT participants
5 Cursor participants
10 weeks Proposed pilot
01Team AI proposal

Executive recommendation

Buy two managed layers—not one universal tool

They solve different parts of the workday and should be measured differently.

Company-wide layer

ChatGPT Business

Research, analysis, writing, connected company knowledge, and repeatable knowledge-work agents.

  • Frontier reasoning with Sol and Sol Pro
  • Company knowledge and managed integrations
  • No training on workspace data by default
  • Central billing, administration, and usage visibility
For knowledge workers already using AI chat
Software-delivery layer

Cursor Teams

An AI-native coding harness for repository context, multi-file agents, verification, and review.

  • Broad frontier-model access and routing
  • Integrated editor, terminal, browser, and subagents
  • Shared rules, skills, privacy, and analytics
  • Standard and Premium seats in one team
For developers and technical builders
Starting principle Standard by default. Premium only for measured heavy users. Revisit the mix at weeks 4 and 7.
02Recommendation

Why move beyond free chat

The upgrade is operational—not cosmetic

Today: fragmented personal use
CapacityInterruptions and tool limits during real work
ContextAd hoc uploads and personal connections
DataIndividual settings and opt-outs
ControlNo shared policy, analytics, or owner
SpendUntracked personal subscriptions and API keys
Target: managed AI workspace
CapacityHigher limits with right-sized heavy-user seats
ContextPermission-aware company knowledge and repository tools
DataBusiness privacy defaults and enforceable settings
ControlNamed admins, policies, budgets, and usage visibility
SpendCentral billing with measurable cost per accepted task
“The choice is no longer adoption versus no adoption. It is unmanaged use versus managed use.”
03Why paid

Interactive cost model

Right-size capacity by person

Knowledge work

ChatGPT Business

$450

8 Standard × $25 + 2 Premium × $125

Software delivery

Cursor Teams

$280

4 Standard × $40 + 1 Premium × $120

04Seat mix

ChatGPT Business

From answers to governed workflows

DimensionFree personal workspaceChatGPT Business
ModelsLuna default; no Sol reasoning tiersSol Medium–Extra High plus Sol Pro
ResearchLimited Deep Research allowanceExpanded research across web, files, sites, and enabled apps
Company contextNo managed organization knowledge layerPermission-aware company knowledge with citations
RepeatabilityPersonal projects and limited schedulingShared projects, workspace agents, custom GPTs, and advanced schedules
Data handlingTraining sharing enabled by default; user can opt outWorkspace inputs and outputs excluded from training by default
AdministrationNo organization controlsRoles, SSO, analytics, spend controls, and central billing
Genuinely Business-specificCompany knowledge, managed GPT creation, workspace agents, governance, and collaboration.
Frame accuratelyProjects and Deep Research exist on Free; paid value is capacity, collaboration, and control.
05ChatGPT value

Model quality

GPT-5.6 raises the reasoning ceiling

Free now uses Luna. Business adds Sol and Sol Pro; Terra and Luna remain available in Work and Codex.

Historical baseline

GPT-5.5 Instant

Fast conversational model and previous default. No clean apples-to-apples comparison with 5.6.

Fast

GPT-5.6 Luna

Free default. Fastest and lowest-cost member of the 5.6 family.

Balanced

GPT-5.6 Terra

Balanced capability and cost for Work, Codex, and API workflows.

OpenAI-published model-level evaluations

Engineering quality, higher is better

EvaluationGPT-5.5LunaTerraSol
SWE-Bench Pro59.4%62.7%63.4%64.6%
DeepSWE v1.167.0%67.2%69.6%72.7%
Terminal-Bench 2.185.6%84.7%87.4%88.8%

GPT-5.5 in this table is the reasoning model—not Instant. OpenAI has not published a clean ChatGPT Sol Pro versus Sol score.

06GPT quality

Paid ChatGPT capabilities

What people can do beyond ordinary chat

01

Deep Research

Editable plan, public web, files, selected sites, and approved apps with a cited report.

Expanded on Business
02

Agent mode

Browser-and-tool workflows that can act, pause for approval, and complete multi-step tasks.

Unavailable on Free
03

Company knowledge

Answers grounded in connected company sources while respecting each user's permissions.

Business-specific
04

ChatGPT Work

Longer tasks that produce documents, spreadsheets, presentations, reports, and Sites.

Expanded on Business
05

Workspace agents

Create, preview, share, and schedule repeatable agents inside the managed workspace.

Business-specific
06

Managed GPTs

Build and publish reusable internal assistants subject to workspace permissions.

Business-specific
07Paid workflows

Coding-model quality

Frontier models are close; cost is not

CursorBench 3.2
Bar: CursorBench scoreRight column: average cost per task when published
ShortlistGrok 4.6, Opus 5, and GPT-5.6 Sol all belong in the frontier pilot.
Do not over-read the rankingThe 0.8-point top-three spread is not decision-grade. OpenRouter rankings measure usage, not quality.
08Coding models

Coding harness landscape

The harness is part of the product

The model name alone does not determine context quality, tool use, verification, safety, or cost.

ProductModel choiceHarness strengthGovernanceBest fit
Copilot BusinessBroad hosted + BYOKGitHub-native and multi-IDEStrong GitHub controlsLowest rollout friction
Claude Code TeamClaude family onlyTerminal-native, deeply co-designedStrong team tierLong autonomous terminal work
Kilo Teams500+ hosted/BYOK/localOpen, configurable agent modesModerate; Enterprise for controlsModel freedom and local inference
VS Code BYOKCompatible custom endpointsFlexible but quality varies by modelFragmented without central setupExisting provider contracts
Independent cross-check: Artificial Analysis v1.4 scores complete model–harness pairs: Claude Code + Opus 5 at 68, Codex + GPT-5.6 Sol at 65, Grok Build at 64, Kimi Code at 63, and OpenCode + Gemini 3.7 Flash at 60.
09Harness landscape

Steering toward correct results

Good UX is only the visible layer

1

Persistent context

Team, repository, path, and user instructions encode conventions once.

2

Evidence tools

Search, read, terminal, browser, MCP, tests, and CI expose real state.

3

Delegation

Subagents isolate exploration, shell output, browser work, and specialist tasks.

4

Deterministic checks

Hooks enforce formatting, tests, secret scanning, and policy—not just prompt advice.

5

Risk boundaries

Allowlists, approval modes, sandboxes, network controls, and protected paths.

6

Recovery

Checkpoints, Git, cancellation, reviewable diffs, and failed-loop detection.

Operating principle Give the agent read, search, and test evidence before broad write or network rights. Put non-negotiable checks in hooks.
10Steering

10-week rollout

A staged pilot with explicit exits

Base license envelope$730 / month
Week 0

Baseline

Two recurring tasks per participant. Define 20 ChatGPT samples and 10 matched coding-task pairs.

Week 1

Configure

Access, privacy, connectors, models, budgets, training, and named metric owners.

Weeks 2–6

Run

Weekly task logs, admin exports, CI evidence, quality audits, and week 2/4 surveys.

Weeks 4 & 7

Right-size

Premium only above 1.5× Standard median usage or repeated limit pressure.

Weeks 8–10

Decide

Scale only when quality holds, safety gates pass, and validated benefit is at least 3× cost.

11Pilot

Measurement system

Pair perception with system evidence

01

60-second task log

Task, model, baseline minutes, actual minutes, outcome, artifact link, manual repair, limits, and incidents.

02

Matched-task design

Compare similar recurring tasks for the same person before and after—not unlike tasks across different people.

03

System evidence

Usage analytics, credits, first CI run, review comments, PR timestamps, and permission or hook events.

04

Blind quality audit

An independent reviewer checks 20% of sampled deliverables against a task-specific acceptance checklist.

05

Value equation

Validated hours saved × Loaded hourly cost ÷ Seats + overage ≥ 3× to scale
12Measurement

Outcome milestones

What success looks like by week 10

ChatGPT10 people
  • ≥8/10active in 3 of the final 4 weeks
  • ≥60accepted work sessions logged
  • ≥25%median matched-task time saved
  • ≥90%company-knowledge citation correctness
  • ≤20%overage as a share of license cost
  • 0severe policy incidents
Cursor5 people
  • ≥4/5active in 3 of the final 4 weeks
  • ≥25%lower lead time over 20 task pairs
  • ≥80%first full CI run passes
  • ≥70%tasks accepted with <30 min repair
  • ≤$20license cost per accepted task
  • 0severe permission incidents
13Milestones

Pulse surveys

Ask at weeks 2, 4, and 10

Keep recall to the last seven days and every survey under five minutes.

Week 2

Learn

Onboarding friction, first paid features tried, initial output quality, and steering baseline.

Intervene early
Week 4

Calibrate

Repeat usage, perceived quality, time saved, integrations, models, and seat-tier fit.

Right-size and retrain
Week 10

Decide

Sustained value, trust, workflow dependency, willingness to continue, and one concrete example.

Scale, extend, or stop
Survey answers explain “why.” Analytics, task logs, PRs, and CI establish “what happened.” Use both.
14Survey cadence

Survey draft

ChatGPT pulse questions

Weeks 2 · 4 · 10
  1. SessionsHow many work sessions in the last 7 days? 0 · 1–2 · 3–5 · 6–10 · 11+
  2. Task mixWriting, research, analysis, company knowledge, files/data, presentations, coding, other.
  3. Output quality1 unusable → 5 ready to use. What share needed no more than minor edits?
  4. Time savedTotal hours saved and one task with estimated before/after minutes.
  5. IntegrationsWhich company sources did you use, and in how many sessions?
  1. Model useSol Medium, High, Extra High, Sol Pro, automatic/unsure, or none—with session count.
  2. Deep ResearchHow many tasks? Report usefulness from 1–5 and optionally link one report.
  3. Paid featuresWork, Agent mode, shared Projects, scheduled tasks, custom GPTs, or Codex.
  4. FrictionInaccuracy, speed, limits, missing source, policy uncertainty, unclear feature, or no suitable task.
  5. Continue?At week 10: impact if removed, continue yes/no/unsure, and one-sentence reason.
Week 2 ≥7/10 report 2+ sessions Week 4 Quality ≥3.5/5; ≥8/10 use Sol Week 10 Quality ≥4/5; median ≥2 h saved
15ChatGPT survey

Survey draft

Cursor pulse questions

Weeks 2 · 4 · 10
  1. SessionsHow many real repository Agent sessions in the last 7 days?
  2. Task mixBug fix, feature, refactor, tests, debugging, review, docs, or investigation.
  3. Model useAuto, Grok, Composer, Claude, GPT, Gemini, Kimi, other, or unsure.
  4. Output quality1 unusable → 5 accepted as produced.
  5. Manual steering1 none → 5 constant supervision.
  6. Manual repair0 · <15 · 15–30 · 31–60 · >60 minutes after the agent stopped.
  1. CompletionShare of tasks that reached the requested result with tests passing.
  2. Time savedTotal hours saved and one PR/task with before/after minutes.
  3. Feature useTab, terminal, browser, MCP, rules, skills, subagents, cloud agents, Bugbot, checkpoints.
  4. Failure loopsHow often did you revert, abandon, or finish manually?
  5. SafeguardsToo restrictive, about right, too permissive, or unsure—with any unsafe/blocked action.
  6. Continue?At week 10: impact if removed, continue yes/no/unsure, and one-sentence reason.
Week 2 ≥4/5 report 2+ sessions Week 4 Quality ≥3.5; steering ≤3 Week 10 Quality ≥4; steering ≤2.5; median ≥2 h saved
16Cursor survey

Decision framework

Scale only when all five gates hold

AdoptionChatGPT ≥8/10
Cursor ≥4/5

Sustained in three of the final four weeks.

Productivity≥25% faster

Median improvement across matched tasks.

QualityNo regression

CI, citation correctness, and review rework meet targets.

Economics≥3× benefit

Validated labor value divided by seats and overage.

Safety0 severe incidents

All lower-risk issues closed with no repeat breach.

Scale

All mandatory gates pass.

Extend 4 weeks

Directional value with a fixable gap.

Stop or redesign

Quality regression, severe incident, or benefit below 1.5×.

17Decision gates

Governance boundary

Know when the self-serve plan is not enough

Team / Business

Start here

  • Central billing and administration
  • Mixed Standard and Premium seats
  • SSO and privacy enforcement
  • Usage visibility and spend controls
  • Shared rules, integrations, and collaboration
Enterprise trigger

Escalate when required

  • SCIM and automated user lifecycle
  • Audit logs, service accounts, and compliance APIs
  • Fine-grained model, MCP, connector, or network controls
  • Custom retention, data residency, or regulated terms
  • Invoice/PO terms and pooled Cursor usage

Do not buy Enterprise because it sounds safer. Buy it when a documented requirement crosses the self-serve boundary.

18Governance

The ask

Approve a measured
10-week pilot

ChatGPT Business10 users8 Standard · 2 Premium
Cursor Teams5 users4 Standard · 1 Premium
Base envelope$730 / monthBefore overage and tax
Scale gate≥3× benefitWith quality and safety intact

Manage what the team already uses. Measure what it actually changes.

19The ask

Appendix

Primary sources

Verified 26 Aug 2026

All prices shown in USD before tax. Product availability, limits, and model catalogs change frequently; re-check before procurement.

Presentation overview

Jump to a slide