OpenAI 2026 hackathon

The ARC AGI Knight Rises

Increasing Chat GPT Sol’s ARC 3 AGI score 6.5x to 50.37%, offering a framework for fluid intelligence, adaptability, and novel problem-solving with only a 100-line prompt.

Solo project by Andrew Melian · 1 likes · 0 comments

Archive position — measured, not model output

1 like on Devpost

506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #2,066 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

The ARC AGI Knight Rises is a self-reported AI experimentation framework developed by one individual (Andrew Melian), designed to improve performance on the ARC 3 AGI benchmark using a lightweight, zero-shot prompt and a single-agent architecture. It leverages Chat GPT Sol as an experimental partner and includes a structured loop for observation, orientation, action, and reflection.

What changed

The author reports a 6.5x improvement in performance over baseline (from 7.8% to 50.37%) on the ARC 3 AGI benchmark using only a <100 line prompt and minimal infrastructure. The approach is described as a general-purpose solution not tailored specifically for this benchmark.

Single most important open question

Is there evidence of any real-world application or traction beyond the hackathon context, or does this represent an isolated technical achievement with no commercial relevance?

Back to contents

What The Product Actually Is

The description states that the product is a zero-shot, lightweight harness captured in a <100 line prompt, which when used with Chat GPT Sol, increased performance on the ARC 3 AGI benchmark from 7.8% to 50.37%. This framework is said to enable agents to exhibit fluid intelligence, adaptability, and novel problem-solving.

It implements a loop structure based on OODA (observe, orient, decide, act), augmented with memory logging and simulation steps. The system uses Sol as an experimental manager and partner, and includes a process of iterative refinement through experimentation and prompt engineering.

  • Claimed functionality: A general-purpose AI agent framework for benchmark performance.
  • Technical components: Prompt-based architecture using CLI, Python, Markdown, tmux, and Sol.
  • Not evidenced Any actual product, deployment, or customer usage beyond the hackathon.

Back to contents

Positioning & Claim Evolution

The author positions this as a framework for fluid intelligence that can be applied across domains without domain-specific tuning. It is described as a zero-shot solution, implying broad applicability and minimal training requirements.

Key claims:

  • The framework enables agents to show fluid intelligence, adaptability, and novel problem-solving
  • It was developed using AI-as-partner principles rather than full automation
  • It achieved results competitive with or better than existing baselines (e.g., CoT vs. the author’s prompt)

There is no evidence of prior positioning or evolution in claims beyond this single project submission.

Back to contents

Target Customer & ICP

The description does not state a target customer or ideal customer profile (ICP). It describes an individual developer (Andrew Melian) working on AI research and experimentation, likely for academic or hackathon purposes.

  • No evidence of:
    • Specific end-user personas
    • Industry verticals
    • Use cases beyond benchmarking

Back to contents

Business Model & Pricing Evidence

There is no evidence of a business model or pricing strategy. The project is described as a hackathon submission with no indication of monetization, licensing, or commercialization plans.

  • Not evidenced:
    • Revenue streams
    • Pricing models
    • Commercial partnerships
    • Product offerings beyond the prompt itself

Back to contents

Technical & Delivery Signals

The author reports building the solution using:

  • CLI, Python, Markdown, tmux
  • Sol (Ultra and Max) for experimentation management
  • Prompt engineering as core method
  • Iterative refinement over 150+ experiments
  • OODA loop with memory system and simulation steps

Key technical details:

  • Uses a single-agent architecture
  • Implements a structured decision-making loop based on OODA
  • Employs memory logging for reflection and learning
  • Designed to be lightweight (under 100 lines of prompt)

Inferences:

  • The approach is prompt-based, not model-specific
  • It relies heavily on human-in-the-loop refinement
  • The solution is not yet deployed or scalable

Back to contents

Traction & Maturity Signals

There is no evidence of traction, adoption, or maturity beyond the hackathon submission.

  • Not evidenced:
    • Customers or users
    • Revenue or monetization
    • Product usage metrics
    • Post-hackathon development or iteration
    • Real-world deployment

The project is described as a single-person effort, and all performance data comes from internal experimentation.

Back to contents

Competitive Context

The description does not provide information about competitors or the broader competitive landscape. It references:

  • Andrej Karpathy’s autoresearch
  • OpenAI’s Cycle Double Cover Conjecture approach
  • ARC 3 AGI benchmark

However, no comparison to other tools, frameworks, or platforms is made.

  • Not evidenced:
    • Competitor products
    • Market positioning
    • Competitive advantages or differentiation

Back to contents

Key Risks & Red Flags

  • Single-person development: The entire project was built by one person (Andrew Melian), raising questions about scalability and long-term maintenance.
  • No commercialization path: No evidence of a business model, product roadmap, or monetization strategy.
  • Hackathon-only context: The work is tied to a single hackathon submission with no indication of follow-up or real-world application.
  • Performance claims lack external validation: Scores are self-reported and not independently verified.
  • Limited scope: The solution is described as a benchmarking tool, not a general-purpose AI platform.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the intended use case beyond the ARC 3 AGI benchmark?
  2. Are there plans to commercialize or scale this framework?
  3. How does the performance improvement translate into real-world utility?
  4. Is there any plan for productization, deployment, or integration with existing AI platforms?
  5. What are the limitations of this approach when applied outside of the benchmark environment?

Back to contents

Investment/Partnership Verdict

This is a self-reported hackathon project with no evidence of traction, revenue, or commercial viability. The author describes an experimental framework that achieved a significant performance gain on a specific benchmark but does not indicate any product, market, or business development beyond the submission.

  • Confidence level: Low
  • Commercial relevance: Not evidenced
  • Investment potential: Not evident from this description alone
  • Partnership opportunity: Not indicated

This represents an isolated technical achievement with no clear path to commercialization or impact.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.