OpenAI 2026 hackathon

Coxsi Agent

Coxsi Agent transforms vague user intent into executable AI workflows, benchmarks multiple models under the same criteria, and recommends the best context-model combination before deployment.

Solo project by minsuk Kim · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,558 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Coxsi Agent is a self-reported tool that helps users transform vague prompts into executable AI workflows, test them across multiple models under consistent conditions, and recommend a suitable model based on real performance metrics (quality, cost, speed). It operates as an integrated part of the PicSeal website, using a five-stage process to guide users through prompt refinement, model testing, and result comparison. The tool is built with HTML/CSS/JS and integrates with existing AI providers like OpenAI, Anthropic, and Gemini.

What changed

The project description indicates this was developed during a hackathon (OpenAI 2026) and submitted to Devpost. It is described as a working five-stage product, not a demo. The author states that the system was built using Codex and GPT-5.6 for development assistance but that core decisions were made by the team.

Single most important open question

Is there any evidence of actual user adoption or real-world usage beyond the hackathon submission? The description does not mention revenue, customers, or product traction beyond its demonstration in a single repository and live demo.

Back to contents

What The Product Actually Is

The description states that Coxsi Agent:

  • Turns a user’s task or existing prompt into a clear and testable Context.
  • Runs this Context on selected AI models and compares results across token usage, API cost, response time, and blind quality scores.
  • Guides users through a five-stage process involving input, context improvement, model testing, result comparison, and saving.
  • Does not replace made-up values for failed or invalid outputs; instead, it marks such cases as unavailable.
  • Recommends a model based on actual test results and user preferences for quality, cost, and speed.
  • Uses a browser-based client in multiple languages (Korean, English, Japanese, Chinese, Spanish).
  • Integrates with the existing PicSeal login system.
  • Supports saving private or public versions of Contexts, allowing reuse without altering originals.

Evidence All of this is self-reported by the author. No external verification exists.

Back to contents

Positioning & Claim Evolution

The description states that:

  • The product aims to help users improve prompts and choose models based on real-world performance rather than general benchmarks.
  • It distinguishes itself from leaderboard-style tools by focusing on user-defined tasks and preferences.
  • It emphasizes fairness in comparison by ensuring all models receive the same approved Context, reference file, and output format.

Inference The author positions Coxsi Agent as a practical decision-support tool for prompt engineering and model selection, not a general-purpose benchmarking platform. This is an inferred claim based on the stated goals.

Back to contents

Target Customer & ICP

The description does not explicitly define target customers or ideal customer profiles (ICP). It implies that users are individuals working with AI models who want to improve their prompts and select models based on performance metrics relevant to their own use case.

Evidence Not evidenced. The author does not describe specific personas, industries, or user segments.

Back to contents

Business Model & Pricing Evidence

The description does not contain any information about pricing, monetization, or business model. It only describes how the tool works and what it does.

Evidence Not evidenced.

Back to contents

Technical & Delivery Signals

The description states:

  • The client is built with HTML, CSS, JavaScript without a frontend framework.
  • It runs inside the PicSeal website at /coxsi/agent/.
  • Communication happens via signed-in APIs on the same domain.
  • The interface supports five languages and responsive layouts.
  • Model results are returned in a shared format (coxsi-model-result-v1).
  • Quality evaluation is blind, meaning model names are hidden during scoring.
  • Duplicate request protection and state persistence after refresh are implemented.
  • Server-side processes manage context approval, model execution, and saving.

Evidence All technical details are self-reported. No external validation or performance data provided.

Back to contents

Traction & Maturity Signals

The description states:

  • The tool was built during a hackathon (OpenAI 2026).
  • A live demo is available at [https://picseal.com/coxsi/agent/](https://picseal.com/coxsi/agent/).
  • It integrates with the existing PicSeal product.
  • Automated tests cover APIs, workflows, login, result display, translations, builds, and system integration.
  • The project includes support for public/private Context versions and version history.

Evidence No evidence of revenue, customer base, or usage metrics beyond the demo and internal development. No mention of user feedback, retention, or product adoption.

Back to contents

Competitive Context

The description does not provide any information about competitors or competitive positioning. It only states that Coxsi Agent is not a general model leaderboard.

Evidence Not evidenced.

Back to contents

Key Risks & Red Flags

  • No revenue or customer data: The tool has no demonstrated traction, adoption, or monetization.
  • Self-reported only: All claims are unverified and based solely on the author’s account.
  • Limited scope: Built for a hackathon, not intended for production use beyond demo.
  • Single developer team: Only one member listed (minsuk Kim), which may indicate limited scalability or depth of development.
  • No pricing or monetization strategy: No indication of how the product would be sold or funded.

Inference The lack of any commercial evidence raises concerns about viability and readiness for market entry.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual user base or traction beyond the demo?
  2. How does the tool plan to scale beyond its current hackathon prototype?
  3. Is there a clear path to monetization or revenue generation?
  4. Are there any existing partnerships or integrations with AI providers or platforms?
  5. What are the long-term plans for model support and feature expansion?
  6. Has the tool undergone any independent security, performance, or usability testing?

Back to contents

Investment/Partnership Verdict

Not evidenced.

There is no evidence of revenue, customers, or product traction beyond a hackathon submission. The description does not indicate whether Coxsi Agent has moved past prototype stage or has any commercial viability. Any potential investment or partnership value must be inferred from the limited self-reported information and cannot be confirmed without further due diligence.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.