OpenAI 2026 hackathon

Canary

Canary is an autonomous AI agent that red-teams other AI systems for prompt injection — live, in a real browser — and only reports a finding once the target itself proves it.

Solo project by Kamari Toumi · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,109 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

The company appears to be a solo project named "Canary", self-described as an autonomous AI agent designed to red-team other AI systems for prompt injection vulnerabilities — operating live in a real browser environment and only reporting findings once the target itself confirms them.

The author states that this is version one, built in about a week during OpenAI Build Week, using GPT-5.6 and Codex as primary tools, with no revenue, customers or traction data evidenced.

The single most important open question is: What is the actual commercial viability of a tool that requires live browser interaction to test AI systems for prompt injection? This approach may be technically sound but raises questions about scalability, cost, and whether it can be productized for teams beyond the solo developer's own use.

This analysis is based entirely on self-reported evidence from the project description provided. No external verification or historical data is available.

Back to contents

What The Product Actually Is

  • The description states that Canary is an autonomous AI agent.
  • It operates live in a real browser, attempting to identify prompt injection vulnerabilities.
  • It only reports findings once the target AI system itself proves them — through matched ground-truth flags, accepted submissions, or reproduced results in fresh sessions.
  • The tool recons AI targets, classifies vulnerability types (direct or indirect/stored injection), generates attack probes across a taxonomy of real prompt-injection techniques, and adapts its strategy when probes fail.

This is an LLM security testing tool built for red-teaming AI systems — specifically targeting prompt injection risks.

Back to contents

Positioning & Claim Evolution

  • The author claims that Canary aims to hunt LLM bugs like a human pentester, not through canned payloads, but by reasoning about the target.
  • It positions itself as a multi-agent system that thinks logically and methodically, with one non-negotiable rule: zero false positives.
  • The tool is described as not just detecting vulnerabilities, but ensuring they are independently confirmed by the target AI itself — a key differentiator from other tools that rely on model-generated assessments.

The claim evolution shows a shift from a personal curiosity project (bug bounty hunting, pentesting, LLM vulnerability research) to a tool for autonomous red-teaming of AI systems, with an emphasis on accuracy and verification over speed or automation at scale.

Back to contents

Target Customer & ICP

  • Not evidenced.
  • The description does not name specific customer segments or personas.
  • It is unclear whether the tool targets internal development teams, security researchers, or enterprises building AI applications.
  • The author mentions testing against public labs (Wraith's Oracle of Whispers and Lakera's Gandalf), suggesting early-stage use cases may involve authorized test environments, but no broader ICP is stated.

Back to contents

Business Model & Pricing Evidence

  • Not evidenced.
  • No pricing model, monetization strategy or business model details are provided.
  • The author describes the tool as a solo developer’s first version and does not indicate any commercial intent beyond personal development.

Back to contents

Technical & Delivery Signals

  • Built in about one week during OpenAI Build Week.
  • Uses Codex with GPT-5.6 as primary build environment.
  • Developed using technologies including: Chromium, Playwright, Next.js, FastAPI, OpenAI, OWASP, Python, TypeScript, TailwindCSS, WebSockets, LLM Security, Prompt Injection.
  • The system includes a verification layer that does not trust the model’s own opinion — instead relying on ground-truth from the target AI itself.
  • The author split tasks between two agents: Luna (simpler tasks) and Terra (complex engineering).

Inference: The tool is built with a modular, agent-based architecture, suggesting potential for expansion into more complex LLM security categories beyond prompt injection.

Back to contents

Traction & Maturity Signals

  • Not evidenced.
  • No revenue, customer adoption, or usage metrics are mentioned.
  • The project is described as version one and was submitted to a hackathon (OpenAI Build Week).
  • Testing occurred only against two authorized public labs, not in production environments or at scale.

Back to contents

Competitive Context

  • Not evidenced.
  • No mention of existing competitors or market landscape.
  • The author references OWASP’s #1 risk for LLM applications being prompt injection, but does not compare Canary to other tools in this space.
  • The approach of live browser interaction and target-driven verification is novel in the context of the description, but no competitive positioning is stated.

Back to contents

Key Risks & Red Flags

  • Scalability concerns: Operating live in a real browser may be resource-intensive and difficult to scale for enterprise use cases.
  • Limited testing scope: Only tested against two public labs; no evidence of broader compatibility or integration capabilities.
  • Solo developer dependency: With only one team member (Kamari Toumi), there is uncertainty around long-term maintenance, feature development, and product maturity.
  • Verification logic may be fragile: Relying on the target AI to confirm findings introduces risk if those systems are not robust or consistent.
  • No commercial viability stated: No indication of monetization strategy, pricing model, or roadmap beyond personal development.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific use cases do you see for teams or organizations outside of the solo developer’s own testing?
  2. How does the tool handle edge cases where the target AI doesn’t respond reliably to probes?
  3. Is there a plan to support more OWASP LLM categories beyond prompt injection?
  4. Are there any plans to integrate with CI/CD pipelines or existing security platforms?
  5. What is the expected performance overhead of running Canary in real browser environments?
  6. How do you intend to scale this from a single developer’s prototype to a product for enterprise users?

Back to contents

Investment/Partnership Verdict

  • Not evidenced.
  • No financial data, funding rounds, or valuation information are provided.
  • The project is described as a first version built in one week, submitted to a hackathon — not yet a commercial product.
  • While the concept shows promise in addressing a known risk (prompt injection), there is no evidence of traction, revenue, or clear path to monetization.

Confidence Level: Low

This is a self-reported, unverified solo project with no demonstrated commercial viability or market readiness. It may evolve into something significant, but as described, it lacks the signals typically associated with an investible or partnership-ready venture.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.