OpenAI 2026 hackathon

cbox

Run Codex unattended inside a hardened container that treats the agent as untrusted - prompt injection can't escape. A shared project brain resumes work across sessions, limits, and machines.

Solo project by Marek Lauko · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,182 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

cbox is a self-reported tool that runs OpenAI Codex unattended inside a hardened Docker container, designed to treat the agent as untrusted while enabling autonomy. It isolates execution through containerization and enforces strict policies via read-only mounts, socket forwarding, and Git-scoped auditing. The system includes a shared on-disk project brain (LEDGER.md, PROGRESS.md) that resumes work across sessions, machines, and engines.

What changed

The author began with a desire to run Codex unattended without approval prompts but within a secure boundary. This evolved into a tool that not only secures Codex but also maintains continuity of work through persistent project state, and builds itself using the same mechanisms it provides.

Single most important open question

Is cbox actually being used in production by anyone other than its creator, or is this an experimental prototype?

Back to contents

What The Product Actually Is

The description states that cbox:

  • Runs OpenAI Codex unattended inside a hardened Docker container.
  • Treats the agent as untrusted through policy enforcement and isolation.
  • Uses read-only mounts for rules and sensitive state.
  • Forwards SSH signing sockets instead of private keys into containers.
  • Audits delegated calls byte-for-byte, with depth-limited Git-scoped execution.
  • Integrates Codex via codex mcp-server, enabling delegation both ways (from console to Codex and vice versa).
  • Provides a shared on-disk brain (LEDGER.md, PROGRESS.md) that resumes work across sessions.
  • Supports optional domain-allowlist egress proxy for traffic control.
  • Was built using Codex itself through its MCP server.

Inference cbox is a developer tool focused on secure automation of AI agents, particularly around Codex, with an emphasis on safety, continuity, and orchestration.

Back to contents

Positioning & Claim Evolution

The author claims:

  • cbox was initially motivated by the need to avoid clicking "allow" on every action in Codex.
  • It evolved from a simple security boundary into a full development console.
  • The tool is built using its own mechanisms — “Codex as a first-class engine” is not marketing but how it was developed.

Inference The positioning has shifted from a niche security-focused utility to a broader developer workflow tool that integrates AI agent automation with persistent context and safe delegation.

Back to contents

Target Customer & ICP

The description does not explicitly name target customers or personas. However, it implies:

  • Developers who use OpenAI Codex extensively.
  • Users seeking secure, unattended automation of AI tasks.
  • Those working in environments where prompt injection or compromised agents are a concern.
  • Individuals or teams that value continuity and context preservation across sessions.

Inference The primary ICP appears to be advanced developers or DevOps engineers who rely on Codex for coding tasks and require secure, persistent workflows.

Back to contents

Business Model & Pricing Evidence

No evidence of pricing, monetization strategy, or business model is provided in the description. The author does not mention any commercial aspects beyond personal use and development.

Not evidenced

Back to contents

Technical & Delivery Signals

The description states:

  • Built with bash, cdi, claude, codex, dante, gpt-5.6, json, mcp, nvidia, ollama, python, ripgrep, socat, socks, sqlite, supervisor, tinyproxy, tmux, ubuntu.
  • Uses Docker for containment.
  • Implements read-only mounts and socket forwarding.
  • Leverages Git-scoped auditing.
  • Integrates Codex via codex mcp-server.
  • Supports multiple delegate tiers (fast, mid, top) per task.
  • Stores project state in plain Markdown files (LEDGER.md, PROGRESS.md).
  • Includes optional domain allowlist egress proxy.
  • Designed for layered containment.

Inference The technical stack reflects a hybrid of scripting, containerization, and AI agent orchestration. The architecture emphasizes security through isolation and control over agent behavior.

Back to contents

Traction & Maturity Signals

The description indicates:

  • The tool was built by the author alone (team size: 1).
  • It is used daily as a personal driver for real project work.
  • The author develops cbox inside cbox, which led to iterative improvements.
  • The roadmap includes concrete next steps like vendor-neutral session hub, shared sessions, Hermes engine integration, infrastructure access, and code-quality refactor.

Not evidenced There is no mention of external users, customers, revenue, or adoption metrics. No evidence of traction beyond personal usage.

Back to contents

Competitive Context

The author notes:

  • They only discovered competition mid-week during the hackathon submission.
  • The tool was submitted to the OpenAI 2026 hackathon on Devpost.

Not evidenced No information about competitors, market positioning, or competitive landscape is provided.

Back to contents

Key Risks & Red Flags

Key risks and red flags inferred from the description:

  • The entire project is self-reported by one individual; no third-party validation.
  • No evidence of external users or adoption.
  • The tool is described as a daily driver for the author, but there's no indication of broader impact or scalability.
  • The roadmap suggests ongoing development without clear commercial milestones.
  • Reliance on Codex (which may not be widely available or stable) could limit utility.

Back to contents

Diligence Questions To Ask The Founders

  1. What is your actual usage frequency and duration with cbox?
  2. Are there any external users or teams currently using cbox in production?
  3. How do you plan to scale beyond a single developer's workflow?
  4. What are the specific use cases where cbox has been applied outside of personal development?
  5. How does cbox handle edge cases like failed delegations or system crashes?
  6. Are there any known limitations or trade-offs with running Codex in this manner?
  7. Do you have plans to support other AI agents beyond Codex, and how would that affect the architecture?

Back to contents

Investment/Partnership Verdict

Not evidenced

The description provides no information on financials, funding, traction, or commercial viability. It is unclear whether cbox represents a viable product for investment or partnership. The tool appears to be an experimental prototype built by one person, with no evidence of market demand or scalability.

Given the lack of external validation and absence of any commercial indicators, this project should not be considered for investment or partnership unless further evidence emerges demonstrating traction, adoption, or a clear path to monetization.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.