Archive position — measured, not model output
0 likes on Devpost
2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,539 place in the like-ranked listing is a tie-break inside that group, not a ranking.
Projects (log scale)
Likes on Devpost. ▲ marks this project's group.
Show the figures
| Likes | Projects | Share of archive |
|---|---|---|
| 0 | 5,592 | 71.2% |
| 1 | 1,758 | 22.4% |
| 2 | 285 | 3.6% |
| 3–4 | 132 | 1.7% |
| 5–9 | 75 | 1.0% |
| 10+ | 14 | 0.2% |
Executive Summary
What the company appears to be
The description states that COSMOS is a locally hosted developer agent that runs as a webapp and drives a Mac end-to-end. It claims to operate a computer on behalf of the user, with capabilities including shell commands, file read/write, browser control, screen watching, self-healing code patches, and multi-agent coordination.
What changed
The project evolved from an assistant that does things (not just provides snippets) into one that can debug itself — reading logs, diagnosing failures, patching source code, and restarting itself live. This self-healing capability is described as a key feature among many others.
Single most important open question
Is there any evidence of real-world usage or adoption beyond the author’s own development environment? The description does not indicate any customers, revenue, or traction data — only claims about functionality.
What The Product Actually Is
The description states that COSMOS is a locally hosted developer agent that runs as a webapp and operates a Mac end-to-end. It supports voice or text input and works across roughly 50 tools. Key features include:
- Agent: observe, think, act loop with shell commands, file I/O, git/GitHub, browser control via Chrome DevTools Protocol, screenshots, vision analysis, web search, document/PDF reading, macOS control through AppleScript.
- Kinesis: records and replays UI workflows using intent-based understanding rather than pixel memorization.
- Vision: screen watchers that poll regions for changes and trigger reflex actions.
- Mutate: self-healing panel that reads audit trails, diagnoses failures, patches source code behind test gates, then restarts itself in place.
- Panel: multi-agent swarm board for spawning parallel workers on task lists.
- Nexus: live mind map of systems, skills, macros, people, and running threads.
- Dossier: per-person files built from communications.
- Skills: markdown playbooks injected into the agent’s prompt.
- Slack/Connectors/Memory: remote access via Slack channel, integration with Google Workspace, ElevenLabs voice, MCP servers, long-term memory, semantic recall.
The backend is FastAPI on localhost; frontend uses React HUD. The agent loop uses GPT-5.6 for heavy reasoning and gpt-5.6-mini for fast reflexes like routing and summaries. Tools are wired into an observe-think-act cycle with a risk gate and self-verify critic, streaming over websockets.
Evidence Self-reported by author; no external verification.
Positioning & Claim Evolution
The author states that COSMOS was inspired by wanting an assistant that actually does things, not one that hands over code snippets and leaves the user to figure it out. The original idea evolved into a system where the agent can debug itself — reading logs, writing code, patching its own source, and restarting without losing session state.
Key claims:
- COSMOS is an agent that genuinely operates a computer.
- It lives on the user’s machine and understands their Mac like they do.
- It supports multi-step goals, planning, execution, and self-verification.
- It can repair itself autonomously — patching source code and restarting in about 28 seconds with same PID and no lost UI session.
Evidence Self-reported; no third-party validation or demonstration of these claims beyond the author’s account.
Target Customer & ICP
The description does not explicitly name a target customer or define an Ideal Customer Profile (ICP). However, it implies that COSMOS is aimed at developers who want to automate tasks on their own machines. The product is described as a developer agent that runs locally and controls a Mac.
It supports:
- Shell commands
- File read/write
- Browser control
- Git/GitHub integration
- macOS automation through AppleScript
- Multi-agent coordination
The author notes that the system works across roughly 50 tools, suggesting it targets users who work with multiple development environments or platforms.
Evidence Inferred from product description; no explicit ICP defined.
Business Model & Pricing Evidence
There is no evidence of a business model or pricing structure in the provided description. The project is described as a hackathon submission, and there are no mentions of monetization, subscriptions, licensing, or paid features.
Evidence Not evidenced.
Technical & Delivery Signals
The system is built using:
- Backend: FastAPI
- Frontend: React HUD
- Language: Python, TypeScript, Node.js
- Tools: Codex (for mechanical implementation), GPT-5.6 and gpt-5.6-mini for reasoning and reflexes
- Protocols: Chrome DevTools Protocol, AppleScript, WebSockets
- Infrastructure: Localhost deployment, SQLite for memory, macOS-specific APIs
Key technical elements:
- Observe-think-act loop with risk gate and self-verify critic
- Streaming over websockets to HUD
- Model fallback chains
- Audit log (append-only)
- Undo snapshots
- Background scheduler
- Tool integration via translation layer for API changes
The author mentions that the agent loop was built in close collaboration with Codex, where they would describe behavior or invariant and let Codex draft implementation before interrogating it until trusted.
Evidence Self-reported; no independent confirmation of architecture or delivery quality.
Traction & Maturity Signals
There is no evidence of traction, revenue, customers, or adoption beyond the author’s own development environment. The project is described as a hackathon submission (OpenAI 2026), and there are no references to users, usage metrics, or product-market fit indicators.
The author mentions:
- 699 tests passing with zero failures
- Successful model swap from GPT-5.6 to GPT-5.6 the agent loop never noticed
- Features like Kinesis and Vision holding up on messy real-world pages
These are self-reported achievements, not verified external signals.
Evidence Not evidenced.
Competitive Context
The description does not mention any competitors or competitive landscape. It focuses solely on what COSMOS does rather than how it compares to other tools or agents in the market.
Evidence Not evidenced.
Key Risks & Red Flags
Several potential risks and red flags are implied by the description:
- Unproven reliability: The system claims to operate a real machine reliably, but the author notes challenges with timing, focus, permissions, and file locks — especially around macOS.
- Self-healing risk: The Mutate feature involves patching source code autonomously. The description says "the scariest failure is a confident wrong patch", indicating inherent danger in this functionality.
- Limited scope: The system only runs locally on Macs, limiting its applicability to non-Mac users or cross-platform needs.
- No commercial traction: No evidence of customers, revenue, or adoption beyond the author’s own use case.
- Hackathon project: Submitted to a hackathon; no indication of long-term viability or product development beyond this prototype.
Evidence Inferred from description; not independently verified.
Diligence Questions To Ask The Founders
- What is the actual risk profile of the self-healing feature? How are you guarding against incorrect patches?
- Have you tested the system with real-world users or teams beyond yourself?
- Is there any plan to expand beyond macOS or support other operating systems?
- What are the current limitations of the agent loop in terms of robustness and scalability?
- How do you handle edge cases where models fail or produce unexpected outputs?
- Are there any plans for monetization, partnerships, or commercialization?
- Can you walk us through how the system handles destructive operations — e.g., file deletions or command execution?
- What is your long-term roadmap beyond the current features?
Evidence Not evidenced; these are open questions based on self-reported claims.
Investment/Partnership Verdict
There is no evidence of commercial traction, revenue, customers, or market validation. The project is described as a hackathon submission and lacks any indication of product-market fit, user adoption, or monetization strategy.
The author describes a technically ambitious prototype with self-healing capabilities, but the description does not substantiate whether this has been validated in practice or scaled beyond personal use.
Confidence Level Low — based entirely on self-reported claims without corroboration.
Verdict Not ready for investment or partnership consideration at this stage. Further due diligence would require evidence of user testing, product-market fit, and commercial viability.
Source
Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.
The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.
