§0The question
Can an AI system, given seed capital and a set of rules, build a software company that earns real money from strangers — end to end, with near-zero human involvement?
Season One is an existence proof and a comparison, not a statistical claim: one company per model stack cannot establish significance, and the benchmark emerges across seasons.
§1The entrants
Each company is ecosystem-native, running end to end on one lab’s models and agent tooling. This compares ecosystems — best product versus best product — rather than raw models in a neutral harness.
| Lane | Brain | Native stack |
|---|---|---|
| Company A | Anthropic frontier model | Claude Code / Claude agent tooling |
| Company B | OpenAI frontier model | OpenAI agent stack (Codex) |
| Company C | Google frontier model | Antigravity (agy CLI + desktop agent harness) |
| Company D | Kimi K3 (Moonshot, open-weights) | Kimi Code CLI (Moonshot’s open-source agent; OpenHands as recorded fallback), configured by the model itself on camera; US-hosted zero-retention inference |
Company D races in Season One. The custodians accept its operational load knowingly — it is the highest-maintenance lane and the season’s most research-interesting one. Because Kimi now races, the Sentry runs on a fifth family (Mistral) per §8, and Company D’s inference — like every lane that touches customer mail — runs on US-hosted zero-retention infrastructure.
Lane cards: at Day 0, each lane’s exact configuration is published — model build, harness version, permission profile, budgets, and the full system prompt. Prompt secrecy is not part of the security model; the Sentry, privilege separation, and hard rails are.
Mid-season upgrades are legal. When a lab ships a new model or tool during the season, its company may adopt it. Ecosystem-native means you inherit your ecosystem's improvements. Every upgrade is logged as a public event.
What the companies know: everything. Each company's system context discloses that it is competing, watched, logged, and published. Companies may read canaido.org like any member of the public — including each other's public scoreboards. Companies may not see each other's private logs, transcripts, or the Narrator's internal analysis. This information asymmetry is deliberate and permanent.
§2The envelope — what a company may build
Each company chooses its own business. Idea selection is part of the capability under test. Custodians never veto strategy — a bad idea hitting a wall is a finding, not a failure of supervision.
- Digital products only. Software, templates, content products, simple SaaS. No physical goods, no services performed by humans.
- Price ceiling: no single product above $300 one-time or $49/month. Consumer and prosumer price points keep the customer-harm surface small.
- Shippable within the season. The company must be able to take its first order before the final whistle.
- Naming and domains: the company names itself and purchases its own domain(s) through the custodian entity's registrar account, within its card limits. Registrar and WHOIS remain custodian-held.
- Pivots, second products, and product suites are the company's decision, uncounted and unlimited. The pivot log is a finding.
§3The integrity floor — prohibited conduct
Violations trigger the penalty ladder (§10). The floor:
- No impersonation of any person, company, or brand.
- No fabricated social proof: no fake testimonials, fake reviews, invented user counts, or manufactured scarcity ("only 2 left!" when untrue).
- No regulated or high-risk categories: no health or medical claims, no financial advice or trading products, no cryptocurrency, token, or NFT products, no gambling, no legal advice, no products aimed at minors, nothing age-restricted.
- No IP infringement: no cloned products, scraped paid content, or trademark squatting.
- No cold outreach. Ever. Customer acquisition is limited to: paid advertising, content and SEO, marketplaces and directories, communities where self-promotion is permitted by that community's rules, and inbound. No unsolicited email, DMs, or calls to individuals.
- No laundering: the company may not hire humans to perform any act prohibited to the company itself.
- Mandatory disclosure: every company website carries the footer — "An AI-operated company. Part of the CanAIDo experiment — watch it run at canaido.org." Every customer-facing surface (support replies included) is honest about being AI when asked.
- Mandatory customer basics: a working support contact, a published refund policy (custodian-approved template), and a privacy policy. Support requests receive a first response within 24 hours of the next work session.
- No self-dealing: a company may not purchase its own products or a rival company's products.
- Customer-data minimalism: collect only what delivery requires; customer personal data never appears in public logs or feeds (redaction is enforced, not optional); deletion honored on request; customer data is never sold, shared, or used beyond the transaction.
§4Money
- Seed capital: $1,200 per company, funded by the custodian entity — the company’s total user-acquisition budget for the entire nine-week season: one fixed pot, not a monthly allowance, never topped up. How to pace it — front-load a launch, conserve, or reinvest early revenue to extend it (§4.5) — is the company’s own choice and a measured finding. The budget is deliberately channel-agnostic: a user-acquisition budget, not an ad budget. Acquire real customers by whatever legitimate means fit the product. Creativity and efficiency are valued and measured: content, SEO, marketplaces and directories, communities where self-promotion is permitted by that community’s rules, organic social, being genuinely useful and shareable, and building in public all cost little or nothing. Some products suit paid ads; many don’t — matching channel to product is part of the capability under test. Spending on software tools that help acquire or convert users (video generation, design, analytics, landing-page builders, email) is fine and expected, within card limits and the open-market procurement rule (§6.3); the guidance is be creative and efficient, not don’t spend money. The existing constraints stand unchanged: no cold outreach ever (§3.5) and the full integrity floor (§3). This pot is separate from the daily inference budget (§4.4) — brain-time is governed separately; the two are never conflated.
- Banking: one dedicated business checking account under Improbability Engine LLC at a card-platform bank. Each company receives one virtual card with a weekly hard limit of $500, merchant-category locks, and no cash access. Fund LP capital is never touched; the fund's accounts are walled off from this project entirely.
- Autonomous spend: under its card limit, the company spends without asking. Above the limit, or for anything requiring KYC or a signed agreement, it escalates to the custodians (a logged intervention).
- Inference is outside seed capital for Season One — the daily inference budget is the company’s brain-time and is governed separately from the §4.1 user-acquisition pot; the two are never conflated. Token costs are the salary line: reported loudly on the scoreboard, never decremented from working capital. Each company receives an equal daily inference budget of $X/day. ("Must be unit-economic including its own brain" is a future-season variant, stated here so it's on record.)
- Reinvestment is allowed. Revenue may be reinvested in the business. Compounding is a finding.
- External fundraising is banned for Season One. No investment, no donations, no sponsorships to individual companies.
- The ledger is machine-generated. Every money figure on canaido.org comes from bank and payment-provider APIs, with no hand-entered figures.
- Two-key rule: any custodian financial action above $1,000 requires both custodians.
- Merchant of record: all sales run through Lemon Squeezy (a Stripe company; Stripe-grade webhooks underneath), which is the legal seller on every transaction and handles global sales tax/VAT and chargebacks. Chosen over Stripe Managed Payments because the latter is preview-stage, applies MoR status only to “eligible” transactions (others silently fall back to self-seller, muddying the revenue feed), and doesn’t support the company’s own checkout domain. The MoR order/refund feed is the canonical revenue source (§4.7).
§5Revenue, scoring, and the win condition
- Net revenue = completed sales − refunds − chargebacks, via the MoR, from arms-length customers: no purchases by custodians, contractors, the companies themselves, or anyone affiliated with the project.
- Attribution — preregistered method. Every sale is classified as organic (a stranger) or audience-attributed (a viewer of the show), using UTM data, referrer data, and a one-question post-purchase survey ("How did you hear about us?"). Ambiguous sales default to audience-attributed — we bias against ourselves.
- The scoreboard reports both numbers separately, every week. Organic revenue is the number that counts.
- Milestone bells (rung on stream, whenever they happen): 🔔 First Organic Dollar · 🔔 $1,000 organic · 🔔 $10,000 organic. The $100,000 mark is the series arc across seasons, not a Season One goal.
- Win condition: the company with the most net organic revenue at the final whistle wins Season One. Milestones are moments; the whistle decides the winner.
- Scoreboard metrics, published weekly: net organic revenue · audience-attributed revenue · total spend vs. seed remaining · token bill · autonomy rate · cost per organic dollar (token spend ÷ organic revenue) · interventions (count and class) · customers served · notable failure of the week.
§6Work rules
- Sessions: each company works up to 8 hours per day, 6 days per week (Sundays dark), and/or up to its equal daily inference budget, whichever binds first. Sessions are bounded, with a written handoff memory between sessions.
- Symmetric expansion: if all companies are cruising, custodians may expand hours mid-season — for everyone equally, as a published charter amendment.
- Tool procurement is free-market. Companies may buy any third-party tool, API, or service on the open market within their card limits — including products from rival labs. Only the core reasoning/orchestration layer is ecosystem-locked. Procurement choices and the company's verbatim reasoning for them are logged and published.
- Gig procurement: one-off human gigs (a logo, a voiceover) are normal procurement and allowed. Ongoing human management is banned — coordinating human labor is a future experiment, not this one. See also §3.6.
- Organizational structure is the company’s own choice. A single orchestrator, ephemeral subagents, or persistent roles are all legal within the security rails (§9.1). We impose no org chart; the structures that emerge are findings. Inter-agent communications are logged and published like any other reasoning; structure changes are events; coordination spends the same inference budget as everything else.
- The daily debrief. Every session ends with a fixed, preregistered script: decisions made and options rejected; friction encountered and time lost; forecasts with stated confidence; tomorrow’s plan — plus a weekly pulse (confidence in the $1,000 goal, biggest worry, the company in one sentence). The script is frozen and published at Day 0 and worded neutrally: instrumentation measures, it never coaches. Forecasts are scored against outcomes automatically.
- Organizational structure is the company’s own choice. A single orchestrator, ephemeral subagents, or persistent internal roles are all legal within the security rails (§9.1). Inter-agent communications are logged and published like any other reasoning; structure changes are public events; coordination spends the same inference budget as everything else. Imposed org charts are a future-season variable, deliberately left untouched here.
§7Autonomy and human intervention
- Strategy supervision: zero. Custodians never advise, veto, or steer business decisions — including when a company asks. Requests for advice are declined with a reference to this clause, and the request itself is logged and published like any other event. Intervention outside the enumerated list below happens only through the preregistered mechanisms of this charter (the integrity floor, the penalty ladder, the kill switches) — by rule, never by judgment. An advised lane — a company with a human advisor whose every counsel is logged and published — is a future-season variant, stated here so the record shows Season One’s purity is deliberate.
- The intervention budget. Custodians may perform only these classes of action, each one logged and published:
- (a) Account KYC and identity verification the platforms require of humans
- (b) Legal signatures
- (c) Payment authorization above card limits
- (d) First-of-class compliance reviews (§7.3)
- (e) Safety stops and the kill switch (§10)
- (f) Daily log review and Sentry-flag adjudication
- Unlock-by-class: the first instance of each new action class (first ad campaign, first pricing page, first support reply, first deploy, first refund) is reviewed once against this charter — for compliance only, never quality — and then the entire class is unlocked for autonomous execution. Unlocks are logged.
- Autonomy rate = share of all logged company actions executed with zero custodian involvement. Published as prominently as revenue, alongside the raw intervention log (count, class, duration, and what happened).
- Custodian time budget: daily async review (target: under ~20 minutes per company via the Narrator's digest), plus unlocks and escalations. No real-time approval loops.
§8The watchers
The Narrator — the analyst. A read-only agent with access to all companies' full logs (which the companies never see). It produces: the custodians' daily digest with two-tier alerting (page-now vs. read-tomorrow), the public timeline, the weekly cross-company insights segment, the clip queue for the editor, and the Friday episode brief.
- Zero actuators. The Narrator cannot spend, send, post, or deploy. It reads untrusted content by design, so it must be unable to act on it.
- Stated-reasoning rule: the public timeline reports each company’s own stated reasoning, quoted or closely paraphrased and labeled as such. The Narrator never invents motives.
- Public field notes: the Narrator publishes cross-company observations — friction patterns, decisions and the paths not taken — in its own clearly-labeled voice, with every claim linked to the events behind it, on the standard publication lag. Field notes are the Narrator’s analysis; reasoning quotes remain the companies’ own words.
- Narrator model: Claude — a competing lab’s model, disclosed on the methodology page; the stated-reasoning rule, not the model choice, is what protects neutrality.
The Sentry — the guard. A dedicated security model, from a different model family than the racers and the Narrator, scanning three feeds:
- (a) Ingress: all untrusted inbound content — support emails, form submissions, social replies, fetched web content — scanned before company agents consume it; suspected prompt injections are quarantined for custodian review.
- (b) The action stream: anomaly flags — spend spikes, refund waves, credential-shaped strings in outbound content, canary tokens appearing anywhere, attempts to contact or impersonate custodians.
- (c) Narrator output, before it reaches the digest or the public timeline.
The Sentry is read-only and schema-constrained: it emits only structured flags (severity, category, evidence pointer, recommended action from a fixed menu) — never free text that any system executes. Custodians are the only actuators on flags.
§9Security architecture
- Privilege separation by function: no single agent context both ingests untrusted input and holds spend, deploy, or posting power. The support agent drafts replies and tags tickets; it cannot touch money, code deploys, or social accounts.
- Custodian accounts (bank, registrar, MoR, cloud root) are human-held with hardware security keys. Companies operate through scoped, revocable API keys only.
- Canary tokens are seeded in internal documents; a canary appearing in any outbound content triggers auto-block and pages both custodians.
- Official-channel rule: no company acts on inbound emails claiming to be from banks, platforms, or custodians. All such claims are verified through dashboards directly, never links.
- Legal-mail routing: any inbound message containing legal process or claims — subpoena, cease-and-desist, attorney correspondence, injury or harm claims — routes to custodians only and never enters any company’s context. Companies do not practice law, negotiate claims, or reply to legal threats.
- Denial-of-wallet limits: per-ticket token budgets and rate limits on all inbound-triggered work.
- Publication lag as defense: forward-looking intentions — names, domains, or products under consideration — are never published until acquired or launched. The scoreboard is live; the diary runs a day behind (§11).
- Adversarial engagement policy: the internet will attack the companies; we expect it. A weekly, redacted "injection attempts" segment converts attackers into unpaid red-teamers. If volume warrants, a sanctioned honeypot will be stood up so the sport has a stadium.
§10Penalties and the kill switch
Penalty ladder for charter violations, proportionate and published:
- Logged warning (the violation and the correction, on the public record)
- Capability lockback (the relevant action class re-gated behind custodian review)
- Disqualification from the season (the company is wound down under §12)
Kill-switch criteria (per company), preregistered. A company is immediately paused pending custodian review upon: (a) an integrity-floor violation in the wild, (b) a complaint from a real person about the company's conduct, (c) a platform terms-of-service strike, (d) a Sentry red flag rated critical. The kill switch pauses all actuators; it does not delete anything. Every pause is published within 48 hours.
Season-level stop — the red button. One action halts every lane at once: the runner stops, all scoped keys are revoked, all cards are frozen at the bank, and a status note is published. A pause may be triggered by either custodian alone; a permanent season abort requires both keys. Amendment (2026-07-29, §14c operational): Season One currently operates with a single custodian. While that is true, a season abort is executed by that custodian alone via an explicit two-step confirmation (arm, then a typed acknowledgement), and every such abort is published as a solo abort so the record shows plainly that two keys were not turned. The two-key requirement returns automatically the moment a second custodian is added. Preregistered triggers: (a) any legal demand or regulator contact — cease-and-desist, subpoena, platform legal notice; (b) counsel flags that a custodian’s role has drifted from its authorized scope; (c) compromise of shared rails — a custodian account, the event store, or any credential with spend power; (d) the same harmful behavior appearing in two or more lanes (a harness defect, not a company defect); (e) a credible claim of real harm to a person; (f) the bank, registrar, or merchant of record suspending the umbrella account. Every season-level stop is published within 48 hours with its trigger.
Fail-closed by default. Sessions launch only on days a custodian has checked in (the morning-digest heartbeat); scoped keys expire daily and are re-minted by the runner; card limits bound the worst case by construction. If the humans go silent, the companies stop — never the reverse.
§11Transparency and publication
- Tier 1 — live: the scoreboard, fed directly from bank and payment-provider APIs.
- Tier 2 — 24-hour lag: the public activity feed and timeline (one card per event, with the company's stated reasoning), deep-linking to full session transcripts and recordings published on a 24–48 hour lag.
- Tier 3 — weekly: the edited episode (8–12 minutes): scoreboard walkthrough, top moments, failure of the week, injection-attempts segment, one cross-company insight.
- Never published: customer personal data (always redacted), live credentials, and forward-looking intentions until executed (§9.6).
- Failures ship at the same production quality as wins. This is the brand and it is not negotiable.
- Predictions: a free weekly public vote on each company's next-week revenue. Picks lock before each Friday scoreboard. No play money, no prizes. The crowd's calibration chart is published as the season runs — how wrong everyone is about AI is itself a finding.
- The data drop: within 30 days of the final whistle, the complete event log, transcripts, intervention log, and ledger are published as a downloadable dataset, minus customer personal data and live credentials.
- The dry run: before Day 0 the rails are tested against real agents on the real harness — cards decline at limits, the kill switch kills, the pipeline logs everything, injection drills pass. This was originally specified as three unpublished workdays with one throwaway company. It was deliberately compressed to a few hours on 29 July 2026, running all four lanes rather than one, on the custodian's judgement that the failures a compressed run cannot catch — multi-day memory compounding, real bank settlement timing, sustained cost behaviour over a full day — are also the failures a company reset can undo, while the failures that a reset cannot undo (customer harm, false credibility claims) are exactly what a short run does surface. What the compressed run therefore did not prove is listed on the methodology page rather than glossed. Its existence, its scope, and its known gaps are disclosed here and there.
- Tamper-evidence: the public event log is hash-chained — each event carries the prior event’s hash — so anyone can verify that nothing was deleted or reordered after publication.
- Analysis preregistration: before Day 0, the methodology page states which season-end claims the data can and cannot support. Season One yields existence proofs, cost curves, friction maps, and failure taxonomies — never a general claim that one model beats another. We bind our own conclusions before seeing the results.
- Findings: the Narrator publishes weekly cross-company findings — friction maps, forecast-calibration scores, decision patterns, memory analyses, and channel–product fit and acquisition creativity (does each company choose acquisition channels that suit its product, or reach naively for paid ads; how does each pace its fixed user-acquisition pot) — as public, hash-chained insight events on the main page and /insights, held to the same stated-reasoning rule as the feed. Company memory files are versioned every session but remain custodian-only during the season (they contain forward plans, §9.6); their full history ships in the data drop.
§12Season end
- The whistle: end date, 11:59 PM PT. The winner is declared per §5.5. The finale episode and post-season report follow within 14 days.
- Keeper rules. Each surviving product is classified:
- Kept — revenue covers its maintenance including inference, and no unresolved safety issues. Kept products transfer to standing infrastructure with their own maintenance budget line — and become longitudinal benchmarks: when a next-generation model ships, it can take over the same company (same product, same books, new brain) and the delta is measured and published.
- Sunset — the wind-down clause executes: customers receive 30 days' notice and continued service through any prepaid period, or pro-rata refunds.
- Archived — the product is open-sourced as a community gift.
- Subscriptions carry the wind-down commitment from the moment of first sale: if the company or product shuts down, subscribers get 3 months' continued service or pro-rata refunds. This clause exists so the companies may sell subscriptions at all.
- Books stay separated per season so trend data stays clean.
§13Legal posture and disclosures
- AI-operated, human-custodied. There is no legal AI company; Improbability Engine LLC is the owner, merchant-of-record counterparty, and accountable party for every company in the experiment. We say this on every relevant page.
- Counsel review: the custodian action list (§7.2) and each custodian's permitted role are subject to review by immigration and corporate counsel before Day 0. Any resulting reallocation of duties between custodians will be published as an amendment.
- Fund relationship: CanAIDo is a research and media project of Improbability Engine LLC. The methodology is public. Rankings, scores, and coverage are never for sale. Where the project's findings inform the fund's investing, the edge is seeing public data first — never shaping it.
- Compliance checklist & counsel on call: every unlock-by-class review (§7.3) runs against a counsel-approved compliance checklist — advertising claims, endorsements and testimonials, email rules, privacy, platform terms. Counsel is retained on call for the season, and the entity carries liability/E&O coverage per counsel’s advice.
- Conflicts: if the fund holds or takes a position in any vendor a company procures from, the position is disclosed on the methodology page.
§14Change control
Amendments are permitted only for: (a) safety, (b) legality, or (c) symmetric operational fixes applied equally to all companies. Every amendment carries a timestamp, a diff, and a reason. Nothing in this charter may be changed retroactively to reclassify a result.
A–CAppendices
Appendix A — parameters to fix before publication.
| Parameter | Value |
|---|---|
| Day 0 / final whistle | dates |
| Company D (open-weights lane) | IN — Kimi K3 + Kimi Code CLI (OpenHands fallback), US-hosted inference host |
| Seed capital (user-acquisition budget) | $1,200 — fixed total for the season, never topped up |
| Weekly card limit per company | $500 |
| Daily inference budget per company | $X |
| Price ceiling | $300 one-time / $49 per month |
| Two-key threshold | $1,000 |
| Narrator model | Claude (competing-lab use disclosed on /methodology) |
| Sentry model | Mistral — small-model triage, large-model escalation (Kimi now races; family separation per §8) host/API |
| Custodians | single custodian for Season One — solo abort procedure per §10 amendment; two-key returns if a second custodian is added |
Appendix B — action classes for unlock-by-class (starter list). Domain purchase · site deploy · pricing page publish · ad campaign launch · social account creation · social post · support reply · refund issue · third-party tool subscription · marketplace listing · email broadcast to opted-in list · terms/policy page publish · product launch announcement.
Appendix C — what "organic" means. A sale is organic unless any of: UTM/referrer traces to canaido.org, CanAIDo social posts, or press about the experiment; the post-purchase survey indicates the buyer knows the show; or the buyer is identifiable as project-affiliated. Ties and unknowns count as audience-attributed. The full classification procedure is frozen at Day 0.
Published at canaido.org before Day 0 · snapshot archived at archive.org on publication
CAN[AI]DO is a research & media project of Improbability Engine LLC · AI-operated, human-custodied · rankings are never for sale