OpenAI 2026 hackathon

roma just talk

STT dictation zero intent to input-execution delay. Talk first. Activate as you continue.

Solo project by Felix celeste · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #6,456 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Roma Just Talk (RJT) is a macOS dictation application built by one developer, Felix Celeste, that aims to reduce the delay between thinking and input by allowing users to begin speaking before activating the dictation tool. The app captures a short rolling audio buffer locally and transcribes both pre-activation and live speech when the user presses a hotkey.

What changed

The author describes an evolution from traditional dictation workflows where users must find a hotkey, wait for activation, then speak — to a model where the user speaks first, activates later, overlapping the two actions. This is presented as a novel interaction design that removes human activation delay from the critical path.

The single most important open question

Is this interaction design actually usable and reliable in real-world conditions, or does it fail due to technical limitations like audio continuity, application compatibility, or transcription accuracy?

Back to contents

What The Product Actually Is

The description states that Roma Just Talk is a macOS dictation app. It captures a short rolling audio buffer locally and transcribes both pre-activation and live speech when the user presses a hotkey.

  • Evidenced from: “Roma Just Talk is a macOS dictation app that captures a short rolling audio buffer locally.”
  • Inferred The app uses Swift for macOS development.
    • Inference based on: “The application also handles: ... built in Swift for macOS and evolved from the VoiceInk codebase.”

Back to contents

Positioning & Claim Evolution

The author positions RJT as an experiment in removing activation from the critical path of dictation. The core claim is that traditional dictation tools force a sequence of actions (stop thought → find hotkey → wait → speak) which interrupts the flow of thinking.

  • Evidenced from: “Most dictation tools force the same sequence: ... The delay is small, but it happens every time. More importantly, it interrupts the moment when the thought is already ready.”
  • Inferred The app aims to make dictation feel more natural by overlapping activation with speech.
    • Inference based on: “Instead of waiting for the microphone before speaking, the user begins talking naturally and presses the hotkey while continuing the sentence.”

Back to contents

Target Customer & ICP

The description does not explicitly name target customers or define an ideal customer profile (ICP). However, it implies a user base interested in efficient dictation workflows, particularly those who type manually or use existing dictation tools.

  • Evidenced from: No explicit mention of specific personas or customer segments.
  • Inferred Likely appeals to macOS users who value speed and efficiency in typing, especially those doing frequent dictation tasks.
    • Inference based on: “Up to 4x my normal manual typing throughput in suitable dictation tasks” suggests a focus on productivity-oriented users.

Back to contents

Business Model & Pricing Evidence

There is no evidence of pricing or business model in the description. The project appears to be a personal experiment or hackathon submission with no indication of monetization strategy.

  • Evidenced from: “This project was submitted to the OpenAI 2026 hackathon on Devpost.” and “No revenue, customer or traction data is available beyond what they state.”

Back to contents

Technical & Delivery Signals

The app is built in Swift for macOS and leverages several technical components including:

  • Global hotkeys
  • Microphone capture
  • Active-application text insertion
  • macOS Accessibility and input permissions
  • Local and cloud speech-to-text paths
  • Application-aware behavior and transcription context
  • Personal vocabulary and custom transcription endpoints

It also includes automated testing and release builds.

  • Evidenced from: “Roma Just Talk is built in Swift for macOS...” and list of features.
  • Inferred The app supports both local and cloud transcription workflows.
    • Inference based on: “Local and cloud speech-to-text paths”

Back to contents

Traction & Maturity Signals

There are no signs of traction, customers, or revenue. The project is described as a single-person effort with no external validation.

  • Evidenced from: “Team size: 1” and “No revenue, customer or traction data is available beyond what they state.”

Back to contents

Competitive Context

The description does not mention competitors or how RJT compares to existing dictation tools. It focuses on the interaction design rather than market positioning.

  • Evidenced from: No mention of competing products or market analysis.
  • Inferred The app likely competes with standard macOS dictation, third-party dictation apps, and speech-to-text services.
    • Inference based on: General nature of dictation tools in the market.

Back to contents

Key Risks & Red Flags

Several risks are implied by the author's own account:

  • Reliability issues: The author notes that reliability is a major challenge, with race conditions, permission problems, and application-specific behavior.
    • Evidenced from: “The core feature already exists. The current limitation is not a missing feature set; it is reliability.”
  • Audio continuity problems: Early versions struggled with joining pre-activation and live audio seamlessly.
    • Evidenced from: “The buffer-to-live hinge... Early implementations treated pre-activation audio and post-activation audio as separate pieces.”
  • User adoption difficulty: The interaction model may be hard to explain or adopt without demonstration.
    • Evidenced from: “A novel interaction can be difficult to explain in text.”

Back to contents

Diligence Questions To Ask The Founders

  1. How does the app handle edge cases like interrupted speech, background noise, or microphone failures?
  2. What is the current level of reliability in real-world usage? Have there been any reported issues during extended use?
  3. Is there a plan to support iOS or other platforms beyond macOS?
  4. How does the app manage privacy and data handling, especially with local audio buffering?
  5. Are there any known compatibility issues with popular applications (e.g., Slack, VS Code)?
  6. What are the performance trade-offs between local vs. cloud transcription?

Back to contents

Investment/Partnership Verdict

Not evidenced.

The project is described as a personal experiment or hackathon submission with no evidence of traction, revenue, or commercial viability. It lacks any indication of a scalable business model or strategic fit for investment or partnership.

  • Evidenced from: “No revenue, customer or traction data is available beyond what they state.” and “This project was submitted to the OpenAI 2026 hackathon on Devpost.”

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.