OpenAI 2026 hackathon

Agentic QA Harness for Governed Tool-Using Agents

AgentOps Receipt proves AI agents use the right tools, match schemas, respect policy, gather evidence, and fail safely before production tool access.

Solo project by Thorsten Zoerner · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #2,411 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

The description states that "Agentic QA Harness for Governed Tool-Using Agents" is a blackbox testing system for AI agents that use tools, APIs, and multi-turn context. The author describes it as a QA harness designed to evaluate agent behavior across multiple dimensions including routing correctness, context handling, input validation, response safety, and governance compliance.

The project appears to be an early-stage technical prototype built by one person (Thorsten Zoerner) for the OpenAI 2026 hackathon. It is described as a sidecar architecture that tests agents from the outside using realistic multi-turn scenarios and generates structured findings with actionable fix prompts.

Key commercial due-diligence questions include: What is the actual product-market fit? How does this relate to existing QA or governance solutions in AI agent development? Is there evidence of adoption or interest from developers or enterprises?

The single most important open question

Does this represent a viable product that addresses a real market need, or is it an experimental prototype with unclear commercial viability?

Back to contents

What The Product Actually Is

The description states that the project is "Agentic QA Harness for Governed Tool-Using Agents" — a blackbox testing system for AI agents that use tools, APIs, and multi-turn context.

It is described as:

  • A QA harness that tests agents from the outside, like a real user would
  • A system that evaluates agent behavior across multiple dimensions including routing correctness, context handling, input validation, response safety, and governance compliance
  • A blackbox testing system that does not inspect private implementation details during scenario execution
  • A system that captures execution traces including request payloads, response payloads, HTTP status, latency, routing metadata, selected capability, execution status, context behavior, and visible answer quality

The harness is described as generating structured findings with:

  • Scenario evidence
  • Suspected root cause
  • Severity
  • Affected capability
  • Recommended fix type
  • Ready-to-use follow-up prompt for a coding agent

Back to contents

Positioning & Claim Evolution

The description states that the project was inspired by the gap in testing tool-using agents, which are different from deterministic software because they route requests, select capabilities, ask follow-up questions, preserve or purge context, call external services, and synthesize answers in natural language.

The positioning claim is:

  • The system tests "the whole agentic behavior, not just the final text response"
  • It evaluates more than "did the answer sound good?" but checks whether agents:
    • Selected the right capability or tool
    • Asked for missing required inputs instead of guessing
    • Preserved useful context across turns
    • Replaced or purged stale context when user changed direction
    • Avoided leaking internal error codes
    • Produced user-facing responses instead of raw execution failures
    • Respected consultation vs. execution mode
    • Followed governance rules around safe tool use

The claim evolution shows a shift from traditional QA (which works well for deterministic software) to testing "agentic behavior" that includes not just language generation but also tool routing, context management, and policy compliance.

Back to contents

Target Customer & ICP

The description states that the system is designed for "governed tool-using agents — agents that operate in domains where transparency, reproducibility, and human oversight matter."

It appears to target:

  • Developers building AI agents that use tools
  • Teams working with "governed" or regulated domains where safety and compliance are important
  • Organizations using AI agents in operational workflows

The description mentions personas-based scenario suites for different user types, such as "Asset Management", suggesting it targets specific professional roles or domains.

However, there is no evidence of specific customer segments, buyer personas, or target industries beyond the example of solar assets in a city.

Back to contents

Business Model & Pricing Evidence

Not evidenced. The description does not contain any information about pricing models, revenue streams, or business model details.

Back to contents

Technical & Delivery Signals

The description states that the system was built around a "blackbox QA sidecar architecture" where:

  • The tested agent runs as a normal service
  • The QA harness interacts through the same HTTP/API surface a real client would use
  • It includes persona-based scenario suites for different user types
  • Multi-turn chat fixtures written as structured test cases
  • Expected routing and capability assertions
  • Context mutation checks for remembering, replacing, and forgetting information
  • Response quality checks for user-facing clarity
  • Failure classification using a structured schema
  • Fix-prompt generation for downstream AI coding agents

The system is described as treating agent quality as an "execution trace, not only as a language-generation problem" and capturing detailed metadata including:

  • Request payload
  • Response payload
  • HTTP status
  • Latency
  • Routing metadata
  • Selected capability
  • Execution status
  • Context behavior
  • Visible answer quality

Back to contents

Traction & Maturity Signals

Not evidenced. The description does not contain any information about revenue, customers, usage metrics, or adoption data.

The project is described as being built by one person (Thorsten Zoerner) for a hackathon, with no mention of any production deployments, customer feedback, or market traction.

Back to contents

Competitive Context

Not evidenced. The description does not contain any information about competitors, existing solutions in the market, or competitive positioning.

Back to contents

Key Risks & Red Flags

Inferences based on self-reported information:

  1. Single-person development: The project is described as being built by one person (Thorsten Zoerner) for a hackathon, which suggests limited resources and potentially unproven scalability.
  1. Prototype nature: It's explicitly described as a hackathon submission with no evidence of production deployment or market adoption.
  1. Unclear commercial viability: The description does not indicate any clear path to monetization or customer acquisition beyond the author's own use case.
  1. Technical complexity vs. execution risk: While the technical approach is detailed, there's no evidence that such a complex system has been successfully implemented at scale.
  1. Market timing uncertainty: The described solution addresses an emerging problem in AI agent development, but there's no indication of market readiness or demand.
  1. Dependency on agent frameworks: The system appears to be built around specific technologies (Node.js, OpenAI, etc.), which could limit its adoption if those frameworks change or become obsolete.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific problems in AI agent testing are you solving that existing QA tools don't address?
  2. How do you plan to scale beyond the current hackathon prototype?
  3. What is your go-to-market strategy for reaching potential customers?
  4. Have you identified any early adopters or pilot customers who might be interested in this solution?
  5. What are the key technical challenges you've encountered during development that could impact delivery?
  6. How do you plan to monetize this product, and what pricing model do you envision?
  7. What is your roadmap for expanding support beyond the current agent frameworks and API formats?
  8. How will you ensure that the fix prompts generated by your system are actually actionable for developers?
  9. What metrics or KPIs do you use to measure success of your QA testing?
  10. How do you plan to handle integration with different AI agent platforms and development environments?

Back to contents

Investment/Partnership Verdict

Not evidenced. The description contains no information about funding rounds, valuations, or investment status.

The project is described as a hackathon submission by one person, with no evidence of any commercial traction, revenue, or investor interest. Without additional information about market demand, customer validation, or business model viability, it's not possible to assess whether this represents a viable investment opportunity or partnership target.

The description states that the system was built for the OpenAI 2026 hackathon and is described as an experimental prototype with no evidence of production deployment or commercial adoption.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.