OpenAI 2026 hackathon

VoiceScene

Build 3D scenes just by talking. Whisper transcribes your voice, GPT-5.6 turns it into scene edits, and three.js renders them live - with a visible JSON scene graph showing every change in real time.

Solo project by Abhijna DS · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #7,600 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

Company: VoiceScene

Self-reported purpose: A voice-controlled 3D scene builder that allows users to create and edit 3D environments by speaking commands.

Key technical elements: Uses Whisper for transcription, GPT-5.6 (gpt-5.6-luna) for interpreting commands, three.js for rendering, and a JSON-based scene graph as the source of truth.

Team size: 1 person (Abhijna DS).

Project context: Submitted to the OpenAI 2026 hackathon.

What changed: The project description is a self-reported account of a hackathon submission. It does not indicate any commercial traction, revenue, or customer adoption beyond the author’s own development effort.

Single most important open question: Is there evidence of product-market fit or early user feedback that would suggest a viable commercial opportunity beyond a proof-of-concept?

Back to contents

What The Product Actually Is

The description states that VoiceScene is a tool for building and editing 3D scenes using only voice commands. It uses:

  • Voice input via push-to-talk with MediaRecorder API
  • Transcription via OpenAI's Whisper API
  • Command interpretation via GPT-5.6 (gpt-5.6-luna) with a fixed set of tools: add_object, modify_object, remove_object, set_camera, animate, stop_animate
  • Rendering via three.js, with a diff-renderer that updates only changed elements
  • Scene state management through a JSON scene graph

The system is described as being built in Codex CLI, using GPT-5.6 as the coding agent.

Inference: The product appears to be a prototype or proof-of-concept for voice-driven 3D scene editing, not a production-ready tool.

Back to contents

Positioning & Claim Evolution

The author states that VoiceScene was inspired by the desire to make 3D scene creation feel like "directing rather than programming." It positions itself as an alternative to traditional GUI or code-based workflows.

Claim: The system allows users to describe what they want out loud and see it appear in real time.

Inference: This is a self-described, unverified claim about user experience and workflow innovation.

The author also notes that the system was built with constraints (e.g., fixed tool schema) to improve reliability over raw LLM code generation.

Claim: Voice commands are interpreted using a constrained LLM agent loop.

Inference: This is a self-reported design decision, not validated in practice.

Back to contents

Target Customer & ICP

The description does not identify any specific customer segments or personas. It does not state who would use this tool or how it would be monetized.

Not evidenced: No indication of target user types, industries, or use cases beyond the author’s own development context.

Back to contents

Business Model & Pricing Evidence

There is no mention of pricing, monetization, or business model in the description. The project is presented as a hackathon submission with no commercial intent or structure described.

Not evidenced: No evidence of revenue streams, pricing models, or customer acquisition strategies.

Back to contents

Technical & Delivery Signals

The system uses:

  • Codex CLI for development
  • GPT-5.6 (gpt-5.6-luna) as the coding agent
  • OpenAI Whisper API for transcription
  • three.js for rendering
  • MediaRecorder API for voice input
  • A JSON scene graph as the single source of truth
  • A diff-renderer to update only changed elements

The author describes building the system in stages, including scaffolding, integrating voice input, and implementing a tool-calling loop.

Inference: The technical stack suggests a prototype built for experimentation rather than scalability or production use.

Back to contents

Traction & Maturity Signals

There is no evidence of traction, adoption, or user feedback. The project is described as a hackathon submission with no mention of users, customers, or usage metrics.

Not evidenced: No data on user engagement, retention, or product-market fit.

Back to contents

Competitive Context

The description does not reference any competitors or market positioning beyond the author’s own innovation claims. It does not describe how VoiceScene compares to existing tools for 3D scene creation or voice-controlled interfaces.

Not evidenced: No competitive analysis or market differentiation described.

Back to contents

Key Risks & Red Flags

  • Unverified claims: All descriptions are self-reported and unverified.
  • Prototype nature: The project is a hackathon submission, not a product with traction or commercial viability.
  • No evidence of user feedback or iteration: No indication that the tool was tested or refined beyond the author’s own development.
  • Limited team size: Only one person built it, suggesting limited scalability or business development.
  • Unproven commercial model: No evidence of monetization or customer base.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the intended user persona for VoiceScene?
  2. Have you tested this with real users beyond your own use case?
  3. How do you plan to scale beyond a single developer’s prototype?
  4. What are the technical limitations of using GPT-5.6 in constrained tool mode?
  5. Are there any plans to integrate with existing 3D tools or platforms?

Back to contents

Investment/Partnership Verdict

Not evidenced: No commercial traction, revenue, or customer data is available. The project is described as a hackathon submission with no indication of product-market fit or business viability.

Confidence level: Low — based entirely on self-reported author description with no external validation or evidence of adoption or monetization.

Verdict: This is a prototype with no demonstrated commercial opportunity or traction. It may be an interesting idea for future development, but it does not meet the criteria for due-diligence readiness or investment consideration at this stage.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.