OpenAI 2026 hackathon

Kirinuki Subtitle Studio

Kirinuki Subtitle Studio is a local-first app that turns long videos into captioned clips using Whisper, VAD, FFmpeg, optional Gemini assistance, and human review.

Solo project by ばたけ とまと · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #4,806 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

Kirinuki Subtitle Studio is a local-first application designed to convert long videos into captioned clips using AI tools like Whisper, VAD, FFmpeg, and optional Gemini assistance, with an option for human review. The author describes it as a tool that guides users through a staged workflow from video import to final export, incorporating technical steps such as speech detection, sentence segmentation, and subtitle synchronization after cuts.

The project is self-reported by one individual developer (ばたけ とまと), who also submitted it to the OpenAI 2026 hackathon. It is not evidenced to have any revenue, customers, or traction beyond its own description. The tool appears to be in early development and has no external validation.

The single most important open question

Is there a viable commercial market for this type of local-first, AI-assisted subtitle editing tool, and does the author have a clear path toward product-market fit or monetization?

Back to contents

What The Product Actually Is

The description states that Kirinuki Subtitle Studio is a local-first app that turns long videos into captioned clips. It uses:

  • Whisper for transcription
  • VAD (Voice Activity Detection) for timing
  • FFmpeg for video processing
  • Optional Gemini assistance
  • Human review capability

It also includes a staged workflow, guiding users from import to editing, preview, and export.

The author notes that the tool organizes these capabilities into presets, aiming to reduce user complexity.

Inference The app appears to be a video editing utility focused on subtitle generation and editing, built using open-source or AI tools, with an emphasis on local processing and human-in-the-loop workflows.

Back to contents

Positioning & Claim Evolution

The author describes the tool as:

  • A local-first app, implying no cloud dependency.
  • Designed to turn long videos into captioned clips.
  • Built using Whisper, VAD, FFmpeg, optional Gemini assistance, and human review.
  • A staged workflow that avoids large operations in one go.

The author also states:

“I learned step by step how to detect speech, determine where sentences should be divided, combine Whisper results with VAD timing, and keep subtitles synchronized after video cuts.”

This suggests a technical learning journey, not a product-market fit or commercial positioning.

Claim

The app is built for creators who want to automate parts of subtitle editing while retaining human control.

Inference The positioning is developer-centric, not customer-facing. It appears to be an experimental tool, possibly for personal use or internal development.

Back to contents

Target Customer & ICP

The description does not identify a specific target customer or ICP (Ideal Customer Profile).

The author says:

“I want to continue developing Kirinuki Subtitle Studio while creating videos about the games, creators, and other subjects I enjoy.”

This implies that the tool is being developed for personal use by the creator, not for a broader audience.

Inference The target customer is likely the developer themselves, or possibly a small group of content creators or developers who are interested in similar tools. There is no evidence of external users or market segmentation.

Back to contents

Business Model & Pricing Evidence

There is no evidence of any business model or pricing structure in the description.

The author does not mention:

  • Revenue streams
  • Monetization plans
  • Subscription tiers
  • Licensing models
  • Paid features

Inference The tool appears to be a personal project, possibly with no commercial intent at this stage.

Back to contents

Technical & Delivery Signals

The app is built using:

  • FastAPI
  • Whisper, Whisper.cpp
  • VAD, Silero
  • FFmpeg
  • Gemini, OpenAI
  • Pydantic, JavaScript, HTML, CSS

It uses a staged workflow and supports local processing, which may imply no cloud dependencies.

The author states:

“I learned step by step how to detect speech, determine where sentences should be divided, combine Whisper results with VAD timing, and keep subtitles synchronized after video cuts.”

This suggests that the tool is technically functional, but not necessarily production-ready or scalable.

Inference The app is a technical prototype, likely built for personal use or demonstration. It may not yet be suitable for commercial delivery.

Back to contents

Traction & Maturity Signals

There is no evidence of traction:

  • No customers
  • No revenue
  • No user data
  • No adoption metrics
  • No product usage logs

The author says:

“Every time I use the application, I discover another improvement or a new idea.”

This implies that it’s still in an iterative development phase, not a mature product.

Inference The project is at an early stage of development, likely in a prototype or MVP phase. No evidence of market traction or adoption.

Back to contents

Competitive Context

The description does not mention any competitors.

However, based on the technologies used (Whisper, VAD, FFmpeg), it appears to be in the video subtitle generation and editing space.

Potential categories include:

  • AI-powered transcription tools
  • Video editing software with subtitle support
  • Local-first video processing tools

There is no evidence of direct competitors or market positioning against them.

Inference The tool may be independent of existing commercial solutions, but there is no indication of whether it fills a gap or competes in any way.

Back to contents

Key Risks & Red Flags

  • No traction or revenue: The tool appears to be a personal project with no evidence of adoption.
  • Single developer: The team size is listed as 1, which may limit scalability and product development speed.
  • No commercial model: There is no indication of monetization or business plan.
  • Unproven market demand: No evidence that there’s a market for this specific tool or workflow.
  • Local-first approach: May limit adoption if users expect cloud-based tools or collaboration features.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the intended user base for Kirinuki Subtitle Studio?
  2. Are you planning to monetize this tool, and how?
  3. Have you tested it with other users beyond yourself?
  4. What are your plans for scaling or expanding the tool beyond personal use?
  5. How do you plan to differentiate this from existing tools in the market?
  6. Do you have any feedback or early user data that shows demand or usage?

Back to contents

Investment/Partnership Verdict

Not evidenced

There is no evidence of:

  • Revenue
  • Customers
  • Traction
  • Market validation
  • Product-market fit

The tool appears to be a personal project, possibly for hackathon submission, with no commercial intent or evidence of adoption.

Inference At this stage, the project does not meet criteria for investment or partnership unless there is a clear plan for monetization and market entry. It is likely in an exploratory phase, not a product-ready one.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.