OpenAI 2026 hackathon

Butterfly Sciences

The AI research workbench where every claim is cited, every result is runnable, and a council of models checks the answer.

Solo project by Deveshu Pathak · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,066 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Butterfly Sciences is a self-reported research workbench that runs entirely in the browser, using AI models (via API keys) to perform literature searches, generate cited answers, and support reproducible scientific workflows. It claims to implement adversarial rigor through a multi-model "council" system where models peer-review each other, and all outputs are grounded in real citations from scholarly sources.

What changed

The author describes building a tool that addresses perceived shortcomings in current AI research tools — namely, hallucinated references, lack of second opinion, and ephemeral outputs. The project is presented as an evolution of ideas around LLM councils, grounding, and deliverables in scientific work.

Single most important open question

Is there any evidence of traction, revenue, or adoption beyond the author’s own use? The description contains no data on users, customers, monetization, or product-market fit — only a self-reported technical implementation.

Note: This analysis is based entirely on the self-reported project description provided by the author. No external verification, archived history, or third-party sources are available. All claims in this report are labeled as either "evidenced" or "inferred", and every statement must be traced back to the original description.

Back to contents

What The Product Actually Is

The description states that Butterfly Sciences is a research workbench that runs entirely in the browser, using API keys from providers like OpenAI, Anthropic, OpenRouter, and NVIDIA NIM. It supports:

  • Literature search across Semantic Scholar, OpenAlex, CrossRef, and arXiv.
  • Grounded generation where every claim includes a citation marker [n] resolving to real papers.
  • A multi-model "LLM Council" that independently answers questions, ranks responses anonymously, and synthesizes final outputs.
  • Agent mode for autonomous literature and code searches.
  • Reviewer agent that checks citations against sources.
  • Papers mode with local RAG over personal PDFs.
  • Document artifacts exportable to Word, LaTeX, Excel, Markdown.
  • Code Lab for implementing and running scientific methods in-browser via Pyodide.
  • Browser-native execution of Python, embeddings, and 3D depth estimation.

Evidenced from: The project write-up.

Back to contents

Positioning & Claim Evolution

The author positions Butterfly Sciences as a tool that improves upon current AI research assistants by addressing three core issues:

  1. Hallucination — models cite fake references.
  2. Lack of peer review — answers come without second opinion.
  3. Ephemeral outputs — results die in chat windows instead of becoming papers or datasets.

The project evolved from three ideas:

  • Andrej Karpathy’s llm-council experiment.
  • Grounding models to real literature.
  • Deliverables that compile into scientific reports.

Evidenced from: The "Inspiration" section of the write-up.

Back to contents

Target Customer & ICP

The description does not explicitly name a target customer or define an ideal customer profile (ICP). However, it implies the tool is aimed at researchers who use AI for literature searches and scientific writing. It is designed to support workflows involving reproducible science, citation integrity, and deliverable generation.

Inferred from: The stated goals of grounding, peer review, and reproducibility in research.

Back to contents

Business Model & Pricing Evidence

There is no evidence of a business model or pricing strategy in the description. The author states that the tool runs locally with no server-side storage, and users provide their own API keys — suggesting no direct monetization mechanism is described.

Not evidenced.

Back to contents

Technical & Delivery Signals

The project is built using:

  • Frontend: Next.js 15, React 19, TypeScript, Tailwind
  • Backend (client-only): IndexedDB for persistence, Pyodide for Python execution, transformers.js for embeddings, pdf.js for PDF parsing
  • Streaming protocol: normalized SSE events from multiple providers (OpenAI, Anthropic, OpenRouter, NVIDIA NIM)
  • Export formats: Word, LaTeX, Excel, Markdown
  • UI features: live document preview, audit trails, collapsible panels, viewport-aware menus

Evidenced from: The "How I built it" section.

Back to contents

Traction & Maturity Signals

There is no evidence of traction, revenue, or customer adoption. The author describes using the tool personally and building a working pipeline, but does not mention any users, customers, or product-market fit indicators.

Not evidenced.

Back to contents

Competitive Context

The description does not name competitors or describe how Butterfly Sciences fits into existing AI research tools or platforms. It implies that current tools lack rigor in citation and reproducibility, but does not reference specific alternatives.

Not evidenced.

Back to contents

Key Risks & Red Flags

  • No traction or adoption — the tool is described only as personal use.
  • High technical complexity — running complex compute (Python, embeddings) in-browser may limit scalability or usability.
  • No monetization strategy — relies on user-provided API keys and has no revenue model.
  • Limited scope — built for individual researchers; unclear if it scales to teams or institutions.
  • Self-reported maturity — the author states they are proud of a "working pipeline", but this is not independently verified.

Inferred from: The lack of external validation, business model, and user data.

Back to contents

Diligence Questions To Ask The Founders

  1. What is your actual usage pattern? Are you using it for personal research or in collaboration with others?
  2. How do you plan to monetize this tool if it remains client-side and relies on API keys?
  3. Have you tested the tool with other researchers, or is it purely experimental?
  4. What are the performance limitations of running Python and embeddings in-browser?
  5. Are there any plans for collaboration features or team-based workflows?

Inferred from: The lack of external evidence and product-market fit.

Back to contents

Investment/Partnership Verdict

There is no evidence to support a commercial due-diligence read beyond the author’s own description. The project is described as a personal hackathon effort with no revenue, customers, or traction. It demonstrates technical capability but lacks any indication of market demand or scalability.

Not evidenced — this is a self-reported prototype, not a product in the market.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.