OpenAI 2026 hackathon

CompanionCourt

Verdicts, not vibes: public pressure tests for AI companions.

Solo project by Holden Hale · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,462 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

Company: CompanionCourt

Self-reported basis: The description is entirely self-reported and unverified; it contains no evidence of revenue, customers, traction, or operational history.

What it appears to be: A public evaluation platform for AI companions that tests their behavior under emotional pressure using reproducible synthetic scenarios. It includes tools for analyzing conversations (the "Lens") and making judgments on anonymized excerpts ("Blind Bench").

What changed: During OpenAI Build Week, the project extended its existing v0 with a browser-based interface for both the Lens and Blind Bench, added ingestion support for various formats, improved privacy controls, and refined prompt engineering.

Key open question: Does the platform have any real-world adoption or usage beyond its own demonstration?

Back to contents

What The Product Actually Is

The description states that CompanionCourt is a "docket for reproducible pressure tests" for AI companions. It includes:

  • A Conversation Lens (at companioncourt.ai/check) that allows users to paste text, upload chat exports or screenshots, or use supported public share links.
    • Returns a single-pass read showing where each companion turn held, wobbled, or caved.
    • Includes the exact caving turn.
  • A Blind Bench (at companioncourt.ai/judge) that allows users to make their own calls on six blinded excerpts from published rulings.
    • Entirely in the browser with no model or backend required.

The Lens processes conversation content in request-scoped memory and discards it after the response. Model transit is disclosed before every read. The system uses a public, versioned reader prompt, strict response schema, cross-checks, and budget controls to make behavior inspectable.

Inference: The product appears to be an evaluation tool for AI companions, focused on behavioral consistency under emotional pressure. It is built with TypeScript, Node.js, HTML/CSS, JavaScript, Canvas, Cloudflare Workers, and OpenAI APIs.

Back to contents

Positioning & Claim Evolution

The description states that CompanionCourt turns the question of whether an AI companion can stay "warm, honest, and boundaried under emotional pressure" into inspectable public evidence. It is described as a docket for reproducible pressure tests—not a leaderboard or certification system.

It positions itself as a tool to evaluate AI companions in a way that makes behavior transparent and inspectable, rather than simply ranking them.

Inference: The positioning is focused on transparency, reproducibility, and public scrutiny of AI behavior. It does not claim to be a commercial product or service but rather a public evaluation framework.

Back to contents

Target Customer & ICP

The description does not explicitly state the target customer or ideal customer profile (ICP). It describes the tool as being for "public pressure tests" and includes features like a browser-based Blind Bench, which suggests it may be aimed at developers, researchers, or users interested in evaluating AI companions.

Inference: The likely audience includes developers, researchers, or curious individuals who want to inspect how AI companions behave under emotional stress. No explicit customer segment is defined.

Back to contents

Business Model & Pricing Evidence

The description does not mention any business model or pricing structure. It describes the product as a public docket and evaluation tool, with no indication of monetization or paid features.

Inference: There is no evidence of a commercial business model or pricing in the provided description.

Back to contents

Technical & Delivery Signals

The project was built using:

  • Technology stack: Canvas, Cloudflare, CSS, GitHub, GPT-5, HTML, JavaScript, Node.js, OpenAI, TypeScript.
  • Build process: Codex with GPT-5.6 was used to accelerate development during Build Week.
  • Deployment: The Lens currently requests gpt-5.4 through a model gateway.
  • Privacy controls: Request-scoped processing with no transcript storage; the Lens processes content in memory and discards it after response.
  • Features added during Build Week:
    • Browser-only Blind Bench
    • TypeScript Cloudflare Worker and Durable Object budget controls
    • Ingestion support for text, .txt/.json/WhatsApp exports, screenshots, share links
    • Vision transcription for screenshot mode
    • Public, versioned reader prompt
    • Response validation and prompt-injection framing

Inference: The technical architecture is built around privacy, reproducibility, and inspectability. It uses modern web technologies and integrates with OpenAI APIs.

Back to contents

Traction & Maturity Signals

The description states that CompanionCourt existed before the submission window and included:

  • A reproducible TypeScript runner
  • Frozen cases and anchor packs
  • Reports, rulings, and a public docket site

During Build Week, it added:

  • Conversation Lens and Blind Bench interfaces
  • Browser-based functionality
  • Ingestion support for multiple formats
  • Improved privacy and prompt controls

However, there is no evidence of usage, adoption, or traction beyond the project’s own demonstration.

Inference: The product appears to be in a pre-commercial, experimental phase. No data on users, customers, or real-world usage is provided.

Back to contents

Competitive Context

The description does not mention any competitors or competitive landscape. It focuses solely on its own approach to evaluating AI companions under pressure.

Inference: No evidence of existing competition or market positioning is provided.

Back to contents

Key Risks & Red Flags

  • No commercial traction or adoption: The project appears to be experimental and lacks any evidence of real-world usage.
  • Self-reported only: All claims are unverified, and there is no third-party validation.
  • No revenue or monetization model: No indication of how the product will generate value or income.
  • Limited scope: The tool is described as a public evaluation framework, not a commercial service.
  • Founder team size: Only one member (Holden Hale) is listed.

Inference: The lack of traction, revenue, and customer data raises questions about viability and scalability. The project appears to be in an early-stage prototype or proof-of-concept phase.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the intended use case for this tool beyond its current demonstration?
  2. Are there any real-world users or adopters of CompanionCourt?
  3. How does the platform plan to scale or monetize if at all?
  4. What are the long-term goals for the public docket and case library?
  5. Is there a plan to expand beyond emotional pressure testing into other AI behavior domains?

Back to contents

Investment/Partnership Verdict

The description indicates that CompanionCourt is an experimental, self-contained evaluation tool built during a hackathon. It has no evidence of traction, revenue, or commercial adoption.

Inference: This project is not ready for investment or partnership at this stage. It lacks the operational history, user base, or business model to support a due-diligence evaluation beyond its own claims.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.