OpenAI 2026 hackathon

Find Your Language Role Model

Find the English creator whose way of thinking already sounds like you.

Solo project by Sumin Kim · 1 likes · 0 comments

Archive position — measured, not model output

1 like on Devpost

506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #1,063 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

The project described by the caller is a tool called Find Your Language Role Model (also referred to as Your Ideal Role Model), which attempts to match English learners with creators whose communication style resembles their own, using AI and natural language processing. It is built as a self-contained demo for an OpenAI hackathon.

What changed

The project was developed by one person (Sumin Kim) over a short time frame, likely during the OpenAI 2026 hackathon. The author states that it was built with specific technologies and methods, including Whisper for translation, GPT-5.6 for style analysis, and cosine similarity matching.

Single most important open question

Is there any evidence of real-world usage or traction beyond the demo? The description does not indicate whether this tool has been used by learners or if it is intended to scale into a product with users.

Back to contents

What The Product Actually Is

The description states that the product:

  • Matches English learners with creators whose communication style resembles theirs.
  • Learners record about two minutes in their native language.
  • Groq Whisper translates the recording into English; raw audio is deleted immediately.
  • GPT-5.6 identifies observable style traits from the learner’s transcript.
  • A centered cosine k-NN search matches against a corpus of 140 creators.
  • The system returns top three matches with evidence and a confidence judgment (strong, clear, or partial).
  • It avoids using voice cloning or mass scraping.

Inference This is an experimental tool built for demonstration purposes, not a commercial product. It uses a combination of LLMs, NLP tools, and vector search to provide personalized English learning recommendations based on communication style.

Back to contents

Positioning & Claim Evolution

The description states:

  • The idea is that copying a native speaker’s rhythm or argumentation may feel unnatural.
  • Instead, learners should shadow someone whose way of thinking already sounds like them.
  • This approach makes English feel more like their own voice in another language.

Inference The positioning shifts from generic advice (“shadow a YouTuber”) to a more nuanced idea: matching learners with creators who share their thinking style, not just content or accent. This is a novel framing within the English learning space, though it’s presented as an early-stage experiment.

Back to contents

Target Customer & ICP

The description states:

  • The target user is an English learner.
  • No specific demographic or segment is mentioned beyond that.
  • The tool does not require a script or account.
  • Learners record in their native language and get matched with creators.

Inference The ICP appears to be self-directed English learners who are looking for personalized, style-based guidance. There is no indication of B2B or institutional use, nor any segmentation beyond learner type.

Back to contents

Business Model & Pricing Evidence

Not evidenced.

Explanation

There is no mention of pricing, monetization, or business model in the description. The project is described as a hackathon submission and demo, with no indication of commercial intent or revenue streams.

Back to contents

Technical & Delivery Signals

The description states:

  • Built with: anthropic, claude, cosine-similarity, groq, huggingface, mstyledistance, next.js, pgvector, postgresql, python, railway, sentence-transformers, styledistance, vercel, vercel-ai-gateway, whisper.
  • Uses Groq Whisper for translation.
  • GPT-5.6 is used for identifying style traits and generating evidence.
  • Centered cosine k-NN search with a 140-creator corpus.
  • The system uses deterministic Python supervisor to control workflow.
  • Demo runs end-to-end in ~57 seconds.

Inference The technical stack reflects a modern, LLM-driven NLP pipeline. It includes vector search, embedding models, and orchestration logic. However, the project is described as a demo with no indication of scalability or production deployment.

Back to contents

Traction & Maturity Signals

Not evidenced.

Explanation

There is no evidence of users, customers, or adoption beyond the author’s own account. No data on usage, retention, or feedback is provided. The system is described as a hackathon demo and not yet a product with real-world traction.

Back to contents

Competitive Context

Not evidenced.

Explanation

The description does not mention competitors or similar tools in the English learning space. There is no indication of market analysis or awareness of existing solutions.

Back to contents

Key Risks & Red Flags

  • No commercial traction or user data: The project is described as a demo, with no evidence of real-world usage.
  • Unverified corpus: The 140 creators include 10 human-verified and 130 AI-drafted candidates. There’s no clarity on how the latter are validated or whether they represent a reliable basis for matching.
  • No clear path to monetization: No pricing, business model or revenue plan is described.
  • Limited scope of use case: The system only matches based on communication style and does not support practice or feedback loops.
  • Self-reported validation: The author states that the core hypothesis (style survives translation) was tested, but no external validation or data is provided.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the source of the 140 creators in the corpus? How were they selected and verified?
  2. Has this system been tested with real learners beyond the demo?
  3. What are the limitations of the matching algorithm, and how does it handle edge cases or ambiguous inputs?
  4. Is there any plan to scale beyond the current demo or hackathon version?
  5. Are there any legal or ethical concerns around using creators’ content for matching without explicit permission?

Back to contents

Investment/Partnership Verdict

Not evidenced.

Explanation

There is no evidence of a commercial product, revenue, or traction that would support an investment or partnership decision. The project is described as a hackathon demo with no indication of scalability, monetization, or user adoption. It may be a promising idea in concept but lacks the signal to justify further due diligence at this stage.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.