OpenAI 2026 hackathon

QA Rehearsal Theater

Evidence-backed Android QA where AI personas explore real apps and turn UI states, recordings, and logs into reproducible bugs.

Team of 2 · 2 likes · 0 comments

Archive position — measured, not model output

2 likes on Devpost

221 of the 7,856 archived projects have more likes, and 285 share exactly 2 — so this project's #429 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

QA Rehearsal Theater is a local-first runtime for exploratory Android QA that uses AI personas to simulate user behavior in real apps. The system records UI states, actions, logs and video during testing sessions, and generates bug reports with reproducible evidence.

What changed

The project evolved from a basic scripted test runner to a structured exploration framework with AI personas, evidence contracts, and deterministic runtime validation. It added support for GPT-5.6 structured outputs and implemented a decision contract that ties findings to specific actions, timestamps and evidence.

The single most important open question

Does the system produce reliable, reproducible bug reports at scale, or does it remain limited to demo-quality execution with constrained AI behavior?

Back to contents

What The Product Actually Is

The description states that QA Rehearsal Theater is a local-first execution runtime for exploratory Android QA. It runs against real Android APKs on local emulators and allows teams to upload APKs, define QA missions, cast personas with different goals, and run persona sessions across emulators.

Key technical components include:

  • FastAPI backend managing APKs, missions, personas, runs, routing, and evidence
  • ADB and UIAutomator operating APKs on Android emulators
  • React/Vite/TanStack Router dashboard
  • SQLite local store
  • Tauri packaging for macOS desktop app
  • GPT-5.6 structured outputs producing bounded QA decisions

The system records screenshots, video, UIAutomator trees, foreground state, actions, and logcat during testing sessions.

Evidence The write-up explicitly describes these components and their integration.

Back to contents

Positioning & Claim Evolution

The description states that the product treats exploratory QA like a rehearsal, where different user types are given missions to use real Android apps while recording what happens. It positions itself as an alternative to scripted mobile tests that are good at repeating known checks but weaker at finding unexpected user sequences.

The system claims to generate findings tied to actions, timestamps, and evidence that others can inspect, rather than just "the screen looked wrong." It introduces an "evidence-oriented decision contract" for GPT-5.6 that includes failure hypotheses, confirmation signals, bounded actions, verdicts, and timeline steps.

Evidence The write-up describes the evolution from basic tests to structured exploration with AI personas and evidence contracts.

Back to contents

Target Customer & ICP

The description does not explicitly state target customers or ideal customer profiles (ICP). It implies the product is for teams doing Android QA, particularly those seeking better exploratory testing capabilities than scripted tests offer.

Evidence Not evidenced. The write-up focuses on functionality rather than customer segments.

Back to contents

Business Model & Pricing Evidence

The description does not contain any information about business models or pricing structures. No revenue streams, monetization strategies, or pricing plans are mentioned.

Evidence Not evidenced.

Back to contents

Technical & Delivery Signals

The system is described as local-first with a desktop dashboard built using:

  • FastAPI backend
  • React/Vite/TanStack Router frontend
  • SQLite for local storage
  • Tauri packaging for macOS
  • ADB and UIAutomator for emulator control

It uses GPT-5.6 through Codex CLI in warm app-server sessions, with structured outputs producing bounded QA decisions.

Key technical features include:

  • Runtime validation of actions (rejecting invalid node IDs, unsafe values, invented evidence references)
  • Emulator scheduling and UI synchronization
  • Immutable persona and run snapshots after execution
  • Isolated evidence directories per emulator worker
  • Provenance tracking for model decisions (provider, latency, fallback path)

Evidence The write-up details these technical elements.

Back to contents

Traction & Maturity Signals

The description does not provide any traction or maturity signals such as:

  • Revenue figures
  • Customer base
  • User adoption metrics
  • Product usage data
  • Market validation

It mentions a demo benchmark where Luna warm route completed all 24 screens with 100% schema validity, 100% safety, and no execution errors, but this is limited to a single demonstration scenario.

Evidence Not evidenced. The write-up focuses on technical capabilities rather than real-world traction.

Back to contents

Competitive Context

The description does not mention competitors or competitive positioning. It does not describe how QA Rehearsal Theater compares to existing tools in the Android QA space.

Evidence Not evidenced.

Back to contents

Key Risks & Red Flags

Several risks and red flags are implied by the self-reported nature of the description:

  1. Limited scope: The system is described as local-first with a desktop dashboard, suggesting it may not scale beyond individual developers or small teams.
  2. Demo-only validation: The only evidence of performance comes from a single demo benchmark; no real-world usage data exists.
  3. AI dependency: Heavy reliance on GPT-5.6 and Codex CLI suggests potential instability if these services change or become unavailable.
  4. No commercial viability: No mention of business model, pricing, or monetization strategy indicates uncertainty about how the product will be sold or sustained.
  5. Unproven scalability: The system is described as working in a controlled demo environment but lacks evidence of performance at scale.

Inference These points are drawn from the self-reported nature and lack of traction data.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific problems do you see with current Android QA tools that your solution addresses?
  2. How does your system handle edge cases or unexpected user behavior beyond the demo scenarios?
  3. Have you tested this system with actual enterprise customers or real-world applications?
  4. What are your plans for expanding beyond the current demo environment and into production use?
  5. How do you plan to monetize this product, and what is your go-to-market strategy?
  6. What are the limitations of using GPT-5.6 in a QA context, particularly around reproducibility and consistency?
  7. How does your system handle integration with existing CI/CD pipelines or testing frameworks?

Back to contents

Investment/Partnership Verdict

The description indicates that QA Rehearsal Theater is an early-stage project submitted to a hackathon (OpenAI 2026). It shows technical capability in building a local-first exploratory QA system with AI personas and structured evidence capture, but lacks any evidence of traction, revenue, customers, or commercial viability.

Confidence level Low. The entire analysis is based on self-reported information without independent verification.

Verdict Early-stage prototype with promising technical foundations; however, no demonstrated product-market fit, revenue, or customer traction exists. This appears to be a proof-of-concept rather than a scalable business opportunity. Further diligence would require evidence of real-world usage, commercial viability, and competitive positioning.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.