OpenAI 2026 hackathon

FORGE

AI changed who can build software. FORGE is my attempt to change how we learn to trust it. | A forensic governance runtime for software audits you can reproduce, inspect, and challenge.

Solo project by Anna Tchijova · 1 likes · 0 comments

Archive position — measured, not model output

1 like on Devpost

506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #1,096 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

FORGE is a forensic governance runtime for software audits that emphasizes reproducibility, traceability, and epistemic discipline in engineering systems. It is described as a system built by one developer (Anna Tchijova) with a focus on architectural failures, broken invariants, and epistemic flaws — not generic bug detection.

What changed

The author states that FORGE emerged from her experience maintaining more than thirty deterministic systems and the recurring engineering problem of building software becoming easier while maintaining justified confidence is not. It is presented as an evolution of her own practice rather than a hackathon idea.

Single most important open question

Is there evidence of real-world usage or adoption beyond the author’s own repositories, and does FORGE have any traction in actual engineering workflows?

Note: This analysis is based entirely on self-reported information from the project description. No independent verification or historical data is available. All claims are unverified.

Back to contents

What The Product Actually Is

The description states that FORGE is a forensic governance runtime for software audits. It includes:

  • A deterministic audit loop: Observe → Hypothesize → Falsify → Verify → Seal.
  • Four execution modes:
    • CLI for local, reproducible audits and sealed reports.
    • Python API for embedding into workflows.
    • MCP (Model Communication Protocol) for connecting to agents with explicit tool boundaries.
    • Multi-agent orchestration for comparing external work products against FORGE’s own evidence and contracts.
  • A set of 20 reusable engineering skills that define how the runtime reasons, audits, patches, preserves evidence, and communicates uncertainty.
  • CRONOS, a tamper-evident audit and reasoning trail, exposed through MCP tooling.

It is described as not relying on clever prompts but instead using documented engineering methods. The system supports both human and AI agent use, with a focus on preserving evidence and managing uncertainty.

Claim: FORGE is a forensic governance runtime for software audits.

Evidence: Author's own description.

Inference: It appears to be a tool for auditing software architecture and invariants rather than general bug detection.

Evidence: The author states it focuses on architectural failures, broken invariants, and epistemic flaws.

Back to contents

Positioning & Claim Evolution

The author positions FORGE as a system that changes how we learn to trust AI-generated code. It is not described as a tool for finding bugs but for auditing systems with reproducible, traceable, and defensible processes.

Key claims:

  • AI changed who can build software.
  • FORGE is about making trust reproducible.
  • It applies a scientific workflow to software audits.
  • False positives are architectural signals, not noise.
  • The goal is not to eliminate mistakes but to make them observable before they become irreversible.

The positioning evolves from a personal engineering practice into a tool that can be used by both experienced developers and those learning through AI.

Claim: FORGE is positioned as a forensic governance system for software audits focused on trust, reproducibility, and epistemic discipline.

Evidence: Author’s own description.

Back to contents

Target Customer & ICP

The author states that FORGE is for people working on complex systems — legal, cryptographic, forensic, ML, infrastructure, or security-sensitive. It can make deep review reproducible and traceable for senior developers, and it may teach reusable skills to those learning through AI.

It is not described as targeting general developers or end-users but rather those who care about long-term stewardship of software systems.

Claim: FORGE targets users working on complex, security-sensitive systems.

Evidence: Author’s own description.

Inference: It may appeal to teams doing compliance, forensic analysis, or high-assurance engineering.

Evidence: The focus on determinism, traceability, and auditability.

Back to contents

Business Model & Pricing Evidence

Not evidenced. No information is provided about pricing, monetization, or business model.

Claim: No evidence of business model or pricing.

Evidence: Author’s own description.

Back to contents

Technical & Delivery Signals

The system includes:

  • A CLI for local audits.
  • Python API integration.
  • MCP tooling for agent interaction.
  • Multi-agent orchestration.
  • 20 documented engineering skills.
  • CRONOS for tamper-evident audit trails.
  • Use of deterministic core, SHA-256 sealing, and versioned schema evolution.

It is built with technologies including CSS, HTML, JavaScript, Python, pytest, OpenAI Codex, and others. It supports local execution without internet or tokens in some modes.

Claim: FORGE supports multiple technical interfaces and uses deterministic methods.

Evidence: Author’s own description.

Back to contents

Traction & Maturity Signals

The author states:

  • She has exercised FORGE across dozens of repositories and controlled experiments.
  • The repository includes 46 MB of sealed HTML and JSON audit artifacts in results/.
  • A public archive contains another 764 MB of audit evidence.
  • It found a real infinite loop in its own Integrity Inspector.
  • It was used to review her own work, including ML projects.

However, there is no mention of external adoption, customers, or revenue. The system appears to be used internally by the author and not yet deployed in production environments beyond her own.

Claim: FORGE has been exercised across dozens of repositories and includes audit artifacts.

Evidence: Author’s own description.

Inference: It is mature enough for internal use but lacks external traction.

Evidence: The lack of third-party adoption or usage data.

Back to contents

Competitive Context

Not evidenced. No information is provided about competitors, market positioning, or competitive landscape.

Claim: No evidence of competitive context.

Evidence: Author’s own description.

Back to contents

Key Risks & Red Flags

  • Lack of external validation: The entire system is self-reported and unverified.
  • Single-person development: Only one team member (the author) is mentioned, raising questions about scalability or long-term maintenance.
  • No revenue or customer data: No evidence of monetization or real-world usage beyond the author’s own work.
  • Unclear adoption path: It's unclear how this would be adopted by teams or integrated into existing workflows.
  • Highly technical and niche: The focus on forensic governance, determinism, and invariants may limit its appeal to a narrow audience.

Claim: Risks include lack of external validation, single-person development, no revenue, and unclear adoption.

Evidence: Author’s own description.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific engineering workflows or teams are currently using FORGE?
  2. How does FORGE integrate into existing CI/CD pipelines or development practices?
  3. Has it been tested in production environments beyond the author's own use cases?
  4. Are there any known limitations or edge cases where FORGE fails to perform as expected?
  5. What is the roadmap for expanding beyond the current 20 governance skills?
  6. How does FORGE handle large-scale codebases, and what performance bottlenecks have been observed?
  7. Is there a plan to open-source additional components or tools beyond the core runtime?

Note: These questions are based on the self-reported nature of the description and aim to probe for evidence that is not present.

Back to contents

Investment/Partnership Verdict

Not evidenced. No information is provided about valuation, funding rounds, or investment interest.

Claim: No evidence of investment or partnership status.

Evidence: Author’s own description.

Inference: Given the lack of external adoption, revenue, or traction, early-stage investment or partnership interest would be speculative without further validation.

Evidence: Absence of any such data in the description.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.