OpenAI 2026 hackathon

DataHarvest

The physical data marketplace bridging the simulation-to-reality gap for robotics. AI labs post bounties, everyday people record spatial data, and serverless AI auto-scores the results.

Solo project by SHIVAM GOEL · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,639 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be: DataHarvest is a self-reported two-sided marketplace for physical robotics training data. The platform allows AI labs to post bounties specifying required physical actions, and everyday users to record those actions via an iOS app that captures sensor data (ARKit poses, LiDAR depth maps, IMU data). A serverless AI pipeline auto-evaluates submissions and pays collectors if the data meets requirements.

What changed: The project was built as a hackathon submission for the OpenAI 2026 hackathon. It is not evidenced to have launched or operated beyond this point.

Single most important open question: Is there any evidence of actual traction, revenue, or customer adoption? The description states no such data exists.

Back to contents

What The Product Actually Is

The description states that DataHarvest is a two-sided marketplace for physical robotics training data. It includes:

  • A mobile app (iOS) built with Swift and ARKit to capture sensor data.
  • A backend pipeline using Python, Modal (serverless GPUs), Supabase, and TwelveLabs.
  • AI models for task grading, object tracking, hand pose estimation, and semantic segmentation.
  • A system where labs post bounties, collectors record actions, and AI auto-scores submissions.

Evidence: The author's own write-up.

Inference: This is a self-reported product architecture. No evidence of actual deployment or usage beyond the hackathon.

Back to contents

Positioning & Claim Evolution

The description states that DataHarvest aims to bridge the "simulation-to-reality gap" in robotics by leveraging smartphone sensor data, which it claims lacks physics in web-scraped video.

Claims made:

  • The platform addresses a #1 blocker for 73% of robotics teams: lack of real-world training data.
  • Web-scraped internet video is insufficient due to missing camera trajectories, LiDAR depth maps, and synchronized IMU data.
  • Smartphones are described as the "most powerful sensor suite on the planet."
  • The system removes the human bottleneck in data validation.

Evidence: Self-reported by author.

Inference: These are positioning claims, not verified facts. No evidence of market research or customer validation is provided.

Back to contents

Target Customer & ICP

The description states that DataHarvest targets:

  • AI Research Labs (as bounties posters)
  • Everyday Collectors (as data contributors)

It also mentions that the platform intends to expand to Android users globally, to unlock "millions of new data collectors."

Evidence: Self-reported by author.

Inference: No evidence of actual customer segments or personas is provided. The description does not indicate whether labs or collectors have been identified or engaged.

Back to contents

Business Model & Pricing Evidence

The description states that:

  • Labs post bounties specifying required actions.
  • Collectors are paid if their submissions meet bounty requirements.
  • AI auto-scores the results, removing human bottleneck in validation.

Evidence: Self-reported by author.

Inference: No pricing model or payment structure is described. The business model is implied to be a marketplace with payments from labs to collectors, but no details are provided.

Back to contents

Technical & Delivery Signals

The description states that DataHarvest was built using:

  • Frontend: Next.js 16 and Tailwind CSS
  • Database & Auth: Supabase PostgreSQL and Edge Functions (Deno)
  • Mobile Capture: Native Swift iOS app with ARKit and CoreMotion
  • AI Backend: Custom Python pipeline on Modal (serverless GPUs)
  • Search: TwelveLabs for semantic video search

Evidence: Self-reported by author.

Inference: The technical stack is described in detail, but there is no evidence of production deployment or performance metrics. The description mentions challenges with GPU orchestration and environment syncing, suggesting early-stage development.

Back to contents

Traction & Maturity Signals

The description states that this was a hackathon submission for the OpenAI 2026 hackathon. It does not provide any evidence of:

  • Revenue
  • Customers
  • Users
  • Product adoption
  • Market traction

Evidence: Self-reported by author.

Inference: No traction or maturity signals are evident beyond the hackathon context.

Back to contents

Competitive Context

The description does not mention any competitors or market positioning relative to existing solutions. It does not state whether similar platforms exist in the robotics data space.

Evidence: Not evidenced.

Inference: No competitive analysis is provided, and no evidence of prior market players is given.

Back to contents

Key Risks & Red Flags

  • The project is described as a hackathon submission with no evidence of commercialization or traction.
  • The author states that serverless GPU orchestration was challenging, suggesting technical risks in scaling.
  • There is no evidence of customer validation, pricing, or monetization.
  • No team size beyond one person is mentioned (SHIVAM GOEL).
  • The platform is described as not yet launched.

Evidence: Self-reported by author.

Inference: These are risks inferred from the lack of evidence for any commercial or operational milestones.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual market demand for this type of physical data in robotics?
  2. Have you validated your pricing model with potential labs or collectors?
  3. How do you plan to scale beyond a single developer and a hackathon prototype?
  4. Is there any evidence of early customer interest or pilot programs?
  5. What are the technical limitations of the current architecture that could impact scalability?

Back to contents

Investment/Partnership Verdict

Not evidenced.

The description states this is a hackathon submission with no commercial traction, revenue, or customer data. There is no evidence of a functioning product, market validation, or business model execution beyond the author's self-reporting.

Confidence: Low. The entire analysis is based on unverified self-description. No evidence of any commercial activity exists.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.