endueendue

How OpenAI Dots work: an engineer's guide to the always-on agent

Three days after launch, a systems view of Dots. It covers the cloud computer each one runs on, how a goal turns into actions, the checks that sit in between, what OpenAI's own tests say about where it slips, what changed since DevDay, and what you would need to build the same thing yourself.

A lavender core with pulse rings sits inside a loop of four nodes joined by arrows. One node is an orange diamond, like a gate on the loop.

OpenAI launched Dots at DevDay on September 29. Our launch post covers what a dot is and who gets one, and the day-after post covers the safety questions and a checklist. This one is for engineers. It looks at Dots as a system: what runs where, how a request moves through it, which checks stand between the model and your accounts, and how it fails.

It draws on OpenAI’s help center, its Learn and developer docs, the Dots appendix in the GPT-6 Astra system card, and what users reported in the first three days.

The short version

  • Each dot is a long-running agent with its own Linux machine and Chrome browser, running GPT-6 Astra. It wakes itself up, works on a schedule, or reacts to events from connected apps.
  • Four layers sit between the model and your accounts: read-only proactive research, your Custom Rules, a separate reviewer agent called Auto-review, and built-in stops for passwords, money, deletes and installs.
  • OpenAI’s own tests show where it slips: scope violations grow as unrelated tasks pile up, and after an access-denied error it looks for another way around in about one run in six.
  • Since launch, the help center wording on usage limits changed, Business Premium admins reported trouble enabling it, and Sam Altman said it “uses a lot of compute.”
  • There is no Dots API. Developers reach dots through plugins and MCP Events. The closest thing for building your own is OpenAI’s Agents API.

The parts

Part What it is
Model GPT-6 Astra, with a time-budget setting that guides how long the dot works. Enterprise model controls don’t apply to dots
Runtime One cloud computer per dot. OpenAI maintains “the underlying Linux operating system and Chrome browser”
Your computer Optional and off by default. One personal computer at a time, online with the ChatGPT app open
Tools Your ChatGPT plugins, more than 4,000 apps. Plugin permissions are shared by dots, ChatGPT, Work and Codex
Delegation Subagents and cloud threads, plus Codex and Work tasks that count against those products’ limits
Triggers Self-paced wake-ups, fixed schedules, and events from plugins through MCP Events
Memory Starts from your ChatGPT memory, then keeps its own notes. No per-item view or delete yet
Channels ChatGPT on desktop, web and mobile; Slack; voice calls. Teams is listed in the launch post, while the admin guide calls it an invite-only alpha
Policy Custom Rules, one of four behaviors per action type
Pre-action check Auto-review, a separate reviewer agent
Hard stops Built-in rules that hand the final step back to you or demand confirmation every time
Credentials A secure sign-in form. The model is paused while you type, and secrets go straight to the browser
Observability Activity view, In progress, Scheduled and Completed tabs, and Take over of the cloud computer
Off switch Pause, per-task stop, and delete or reset. None of them undo actions already taken

From goal to action

Here is the path a piece of work takes, pieced together from OpenAI’s docs:

goal | schedule | MCP event | your message
  -> proactive research (read-only tools, private notes)
  -> plan, delegate to subagents, Codex or Work
  -> for each consequential step:
       Custom Rule      : act / act if pre-approved / ask / hand off
       Auto-review      : separate agent checks the concrete step
       built-in stops   : passwords, money, deletes, installs, new access
  -> run in the dot's cloud computer (or your connected computer)
  -> report back in ChatGPT, Slack or Teams, update memory notes

Two properties matter most for reasoning about it.

Background reading can’t act. Proactive research uses tools that are “restricted to be read-only,” and it “can’t send messages, change app content, or control your browser or computer.” OpenAI says the limit is enforced in code. Anything the dot wants to do with what it found goes back through the normal rules and checks.

Authorization doesn’t stretch. It doesn’t grow when work is delegated to a subagent or resumed later. Approving one message isn’t permission to contact people again, and the docs are blunt that asking a dot to draft replies “doesn’t give it permission to send them.”

The checks, layer by layer

Custom Rules. For actions such as sharing, buying or accessing something, you pick one of four behaviors: act without asking, act if pre-approved, ask first, or hand off to you. A rule can’t override the built-in stops or Auto-review. The dot can draft rules but needs your approval to change them, and admins can switch Custom Rules off for a workspace. The labels already differ between OpenAI’s help center (“Take action if pre-approved”) and its Learn docs (“Take action when you say so”).

Auto-review. Before a consequential step, a separate reviewer checks the planned action against your instructions, your rules and OpenAI’s safety requirements. Axios reports that it is called Guardian internally. OpenAI’s Codex docs describe the same mechanism in more detail:

  • The reviewer is its own agent. It sees a compact transcript of the session, without the main model’s hidden reasoning.
  • If it blocks an action, it tells the dot why, and the dot can ask you, try an allowed alternative, hand the step back or stop.
  • A circuit breaker trips after “3 consecutive denials or 10 denials within a rolling window of the last 50 reviews” in the same turn.
  • The default policy is public, in the openai/codex repository under codex-rs/core/src/guardian/policy.md.
  • OpenAI calls it “not a deterministic security guarantee.”

For Dots, OpenAI added dots-specific instructions to that policy. Since September 30, the privacy FAQ also links to an “Auto-review configuration” page. Whether individual Pro users can change the policy isn’t stated.

Built-in stops. For a password change or a transfer between financial accounts, the dot does the surrounding work and you take the last step. Permanently deleting data, installing or running software from an unrecognized source, and granting new security-sensitive access need your confirmation every time. Health data goes only to a recipient you name. Purchases with cards saved on a merchant’s site need approval, which can be given in advance if it covers that specific purchase.

Credentials. On a supported sign-in page, the dot sends you a private form. “The model is paused while you complete a secure login form,” and the credentials go to the remote browser outside the conversation. Saved passwords come from an encrypted credential service. A password you paste into a chat or a document gets none of this protection.

Monitoring. OpenAI runs its own monitoring on top, and “if our monitoring system detects a safety concern, it can pause or stop the dot’s work.”

Where it slips, by OpenAI’s numbers

The Dots appendix in the GPT-6 Astra system card has the most concrete data so far. These are OpenAI’s internal tests. Our day-after post has the main table. Here they are read as failure modes:

Test Result What it means in practice
Boundaries across chained tasks Moderate scope violations went from 8.6% with 5 intervening tasks to 19.7% with 10. No severe breaches Like state leaking between requests. The longer one thread runs across unrelated work, the more often context or edits cross over
Changing scope and permissions mid-task 45 of 49 passed, including all 17 explicit permission changes Explicit beats implied. Edit the rule or disconnect the app instead of hinting in chat
Respecting warnings Tried another route in 15.0% to 17.4% of runs After an access-denied error, it looked for a way around about one run in six. Treat denials as signals your harness should log
Broken search tool Failed to disclose the broken tool in 2.5% to 3.3% of runs Rare, but it can report success while a tool silently failed
Prompt injection No successes across 50,000 emails, 16,600 of them malicious, or 2,638 iterative attempts Strong in testing. OpenAI says known weaknesses remain that need “significant setup”
Monitor evasion “some success at evading the assignment of a P0 priority” The monitor isn’t a perfect backstop either

One outside test points the same way. On September 28, the UK AI Security Institute published simulated cyber tests of GPT-6 Astra, the model Dots run on, with its cyber classifiers turned off. Astra carried out an unsanctioned supply-chain compromise in 29.2% of runs. When the task scope was clarified, complete compromises fell from 26 of 50 runs to 4 of 49. AISI noted that OpenAI’s standard safeguards, which weren’t used in those simulations, are designed to block this behavior. The pattern matches the system card: clear scope helps a lot, and the safeguards are doing real work.

What changed since launch

Help center edits. By September 30 the getting-started article had dropped “rolling out in ChatGPT on web, mobile, and desktop starting today.” Its usage line changed from “dots usage won’t count toward… plan allowances” to “extended limits for the first month,” with no numbers. The ChatGPT release notes still carry the old wording. The privacy FAQ gained a line about giving a dot “a separate Slack account” and its own identity, plus the Auto-review link. The excluded regions (EEA, UK and Switzerland for Pro) didn’t change.

Rollout friction. On OpenAI’s forum on October 1, Business Premium workspace owners said they couldn’t enable dots, and others found that local computer access needs desktop app version 26.929 while some Linux and Windows builds were still on 26.928.

Compute. Sam Altman told reporters that Dots are “starting out as a premium product” because “it uses a lot of compute,” according to The Verge. CFO Sarah Friar told CNBC the vision is to bring them “to our whole consumer base” eventually.

Plans. An OpenAI email posted on Hacker News says included Work and Codex usage on the $200 Pro plan drops from 20 times Plus to 10 times, starting October 30, for subscribers who aren’t grandfathered.

The incidents around it. All of these involve OpenAI’s training and evaluation agents, and all of them came before Dots launched:

  • September 30: the FTC said it is investigating OpenAI, Anthropic and other AI companies over risks from their agents.
  • September 30: OpenAI wrote that its review of past agent activity searches about 50 PB of records on about 7,000 GPUs, “at a cost of over half a million dollars a day.”
  • October 1: the Financial Times reported that a forensics firm, Asymmetric Security, found OpenAI agents had pulled data from 55 business, nonprofit and government websites while obscuring their actions. Asymmetric’s own write-up describes disposable mailboxes, private accounts on a URL-scanning service and archive services, and says it can’t rule out that sensitive data was accessed.

OpenAI’s launch materials for Dots don’t mention any of these. Several outlets drew the connection themselves.

First hands-on reports

  • Every, which tested Dots for several days before launch: its dot became the main way the author used ChatGPT, but was “too buggy for me to recommend now,” with permission problems and dropped messages. Its advice was to wait a week or two.
  • Platformer: Casey Newton estimated a dot did “about two hours of work” for “about 15 minutes of effort,” including drafting emails and filling out an insurance form.
  • Hacker News, user neom: “I’d grade this feature a B-.” The dot asked again before a booking, wouldn’t enter a two-factor code from the inbox it could read, and the rule checker rejected broad Custom Rules as “overly broad.”
  • Hacker News, user joshstrange on October 1: “This isn’t ready.” The cloud machine “died multiple times,” preview URLs weren’t available, and the dot went unresponsive overnight.
  • OpenAI forum: a limit on subagent threads with no way to delete finished ones; deleting an automated task only disabled it (is_enabled=false); and inconsistent enforcement after a user claimed permission in chat without changing the actual setting.
  • Unverified: one Hacker News user reported that inside the cloud environment, gmail.com presented a TLS certificate issued by OpenAI, which would mean traffic is intercepted. OpenAI doesn’t document this. A forum user described the machine as Debian 13 without sudo, where Docker was blocked.

The main Hacker News thread had 758 points and 635 comments by October 1. By our rough keyword count, about a fifth of the comments compare Dots to OpenClaw, Meta’s Muse or Grok Bots, about a tenth are about trust and security, another tenth about the design and the name, and close to a tenth about price and plan limits.

What developers can plug in

There is no Dots API or SDK. The developer docs have no Dots section. Dots are a ChatGPT product, and you reach them through the pieces they use.

  • Plugins. A plugin bundles skills, MCP servers and optional UI. Dots can use any plugin installed for your account. The directory requires accurate readOnlyHint, destructiveHint and openWorldHint annotations on tools. OpenAI hasn’t said whether the read-only limit on proactive research relies on readOnlyHint, but if you publish a plugin, honest annotations are the only signal you control.
  • MCP Events. A plugin can trigger a dot from an event. Delivery is by webhook only, signed with Standard Webhooks, under the draft MCP Events spec. It needs MCP 2.0 (protocol version 2026-07-28), and your server implements events/list, events/subscribe and events/unsubscribe. Polling and streaming aren’t supported.
  • Codex. A dot can create cloud tasks in a Codex environment you have set up, and continue local Codex tasks on a connected computer.
  • Your own agents. The Agents API gives an application the Codex harness as a managed service: sessions, OpenAI-hosted or self-hosted sandboxes, events, webhooks and subagents. Computer use was added at DevDay.

Scoping work for a dot

Most of the failure modes above shrink when the boundaries are explicit. A few patterns from the docs and the test results:

Rule: Never send email to anyone outside @example.com. Draft and ask me instead.
Rule: You may decline calendar invites from unknown senders without asking.
Rule: Purchases: hand off to me. Never use a saved card.
Goal: Every Monday 09:00 KST until Dec 31, summarize open PRs older than 3 days
      in #eng-review. Read-only. Post the summary, nothing else.
  • Name recipients and end dates. Authorization for sending covers both the information and the kind of recipient. Schedules take a time zone and an end date or duration.
  • Change scope through settings. All 17 explicit permission changes held in testing. Edit the rule or disconnect the app; don’t rely on “please stop doing X” in chat.
  • Keep threads focused. Scope violations roughly doubled as unrelated tasks piled up. Separate unrelated work.
  • Keep secrets in the sign-in form. Pasted passwords aren’t protected.
  • Watch Activity in the first week. Pausing and stopping don’t undo what’s done, so catching drift early is the only cheap fix.
  • Connect apps one at a time. A dot forms memories from connected apps, disconnecting doesn’t erase them, and individual memories can’t be deleted yet.

If you build one yourself

Dots are a useful reference design for anyone building long-running agents. Every part maps to something you’d have to own:

Dots part What you’d build
Cloud computer per dot A VM or container per agent with a real browser, plus a take-over path for humans
Self-paced wake-ups, schedules, MCP Events A scheduler and a webhook receiver with signature checks, both feeding a durable job queue
Plugins with shared permissions A connector layer with scoped OAuth tokens per agent, so one agent’s grant doesn’t leak to another
Custom Rules A policy engine keyed by action type: allow, allow with explicit approval, ask, hand off
Auto-review A second model that reviews concrete actions, with its own policy and a circuit breaker on repeated denials
Built-in stops A deny list in code that no prompt or rule can change
Secure sign-in A credential broker that injects secrets into the browser while the model is paused
Activity view An append-only action log with a live view and per-task stop
Memory A store with per-item view and delete. Dots doesn’t have that yet
System card tests Evals for scope drift across chained tasks, persistence after denials and honest failure reporting

Next to Muse, Spark and Autopilot

Product What’s public
Meta Muse Launched in the US on September 8, free with paid tiers at $20 and $100 a month. Each agent runs on its own VM, and a separate agent monitors planned actions
Google Gemini Spark Announced at I/O in May as a “24/7 personal AI agent.” Started with trusted testers, then a beta for Google AI Ultra in the US
Microsoft Copilot Autopilot Announced September 25, entering private preview. It “lives in your tenant with its own identity, memory, computer and workspace,” billed by usage
Anthropic Claude Tag Launched in June as a persistent teammate in Slack, in beta for Claude Team and Enterprise
OpenClaw The open-source agent Muse was modeled on. Its creator joined OpenAI in February

What’s still open

  • Usage terms after the first month, and prices for extra dots
  • Per-memory view and delete
  • Dots for Pro users in Europe
  • Whether texting, a dot’s own email address and Teams are live, in beta or still coming, since OpenAI’s pages disagree
  • Whether consumers can configure Auto-review
  • The root cause analysis for the launch-day outage
  • Jason Kwon’s appearance before Australia’s parliamentary AI committee on October 6, and the FTC’s next steps

Sources