OpenAI 2026 hackathon

PHI Mask

Automatically mask protected information from AI so your personal and private data stays safe.

Solo project by Benson Wong · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #5,924 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

PHI Mask is a self-reported tool that claims to automatically mask protected health information (PHI) and personally identifiable information (PII) from text, images, and PDFs so that users can safely input this data into public AI systems. The author describes it as a compliance-focused solution built for individuals working with sensitive data — such as doctors or lawyers — who want to protect their privacy while using AI.

What changed

The project pivoted from an initial idea focused on token savings in LLMs (based on rtk) to a compliance-oriented tool after discovering issues with the original approach. It now focuses on masking PHI/PII for safe use with public AI, particularly targeting users who work with patient data.

Single most important open question

Is there any evidence of actual user adoption or traction beyond the author’s own development and testing? The description contains no mention of customers, revenue, usage metrics, or market validation — only self-reported claims about functionality and intent.

Back to contents

What The Product Actually Is

The description states that PHI Mask:

  • Finds PHI/PII in images, PDFs, or text.
  • Masks this data so it can be safely used with public AI.
  • Provides a masked version of the content (e.g., “B*** W***”) that allows users to unmask real content later.
  • Works offline using Tesseract OCR and regex-based detection.
  • Uses GPT-5.6-luna for processing large-scale name/address datasets.

It is described as a tool designed to help users protect sensitive data when interacting with AI systems, especially in regulated environments like healthcare or legal work.

Inference This appears to be a proof-of-concept or early-stage prototype built by one developer (Benson Wong), likely for a hackathon. There is no evidence of commercial deployment or integration into existing workflows.

Back to contents

Positioning & Claim Evolution

The author reports:

  • Initially intended to optimize token usage in LLMs but pivoted due to performance degradation concerns.
  • Later shifted focus to compliance and safety, driven by personal need (wife as a doctor).
  • Positions itself as a tool for protecting private data when using AI — particularly relevant for professionals handling PHI.

Claim

PHI Mask is positioned as a privacy-preserving layer that allows safe interaction with public AI systems without exposing sensitive information.

Inference The evolution suggests the team started with a technical optimization goal but found a more compelling use case in compliance and data protection. However, no evidence exists of strategic positioning or market research beyond personal motivation.

Back to contents

Target Customer & ICP

The description states:

  • The tool was built for someone working with patient data (a family medicine doctor).
  • Intended to support low-tech users who are concerned about AI but must adapt to using it.
  • Targets professionals in healthcare and law firms where compliance is critical.

Inference The primary target appears to be individual practitioners or small teams within regulated industries who need to comply with privacy laws while leveraging AI tools. However, no evidence of actual customer segments or personas exists beyond the author’s own experience.

Back to contents

Business Model & Pricing Evidence

Not evidenced.

The description does not contain any information about:

  • Revenue model
  • Pricing strategy
  • Monetization plans
  • Subscription tiers or licensing options

Inference There is no indication that a business model has been developed or tested. The project seems to be in an early prototype phase, possibly intended for internal use or demonstration purposes.

Back to contents

Technical & Delivery Signals

The description states:

  • Built offline using Tesseract OCR.
  • Uses regex patterns and GPT-5.6-luna for processing large datasets.
  • Includes WYSIWYG UI overlay functionality.
  • Struggled with OCR artifact tolerance, multi-cultural name detection, and typo handling.
  • Was developed during a hackathon.

Inference The tool is technically complex, involving OCR, regex matching, and AI-assisted pattern recognition. However, it's unclear whether this has been scaled beyond a prototype or integrated into broader systems. The reliance on GPT-5.6-luna implies some level of automation but also suggests dependency on external tools.

Back to contents

Traction & Maturity Signals

Not evidenced.

There is no mention of:

  • Customers
  • Users
  • Revenue
  • Product adoption
  • Market feedback
  • Iteration history or product maturity

Inference This appears to be a hackathon submission, not a mature product. No evidence supports any form of traction or commercial viability.

Back to contents

Competitive Context

Not evidenced.

The description does not reference:

  • Competitors
  • Existing solutions in the PHI/PII masking space
  • Market size or competitive landscape

Inference No competitive analysis is provided. The author does not discuss how their solution compares to others, nor whether there are similar tools already available.

Back to contents

Key Risks & Red Flags

Key risks and red flags based on self-reported information:

  1. Unverified claims: All features and capabilities are self-reported without independent validation.
  2. Single-person development: Only one team member (Benson Wong) is mentioned, raising questions about scalability or long-term maintenance.
  3. Dependency on proprietary AI models: Reliance on GPT-5.6-luna may pose risks related to availability, cost, and control.
  4. No commercial traction: No evidence of users, customers, or revenue indicates the product has not yet entered the market.
  5. Offline OCR limitations: The description notes challenges with offline OCR accuracy, which could limit usability in real-world applications.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific compliance standards does PHI Mask aim to support (e.g., HIPAA, GDPR)?
  2. Has the tool been tested with actual users from target industries?
  3. Are there plans for monetization or commercial deployment beyond the current prototype?
  4. How is the accuracy of PHI/PII detection validated in practice?
  5. What are the limitations of the offline OCR approach and how do they impact real-world use cases?
  6. Is there any plan to integrate with existing AI platforms or workflows (e.g., LangChain, LlamaIndex)?
  7. What is the roadmap for scaling beyond a single developer?

Back to contents

Investment/Partnership Verdict

Not evidenced.

There is no evidence of:

  • Financials
  • Customer base
  • Product-market fit
  • Strategic partnerships
  • Market traction or growth indicators

Inference At this stage, PHI Mask appears to be an experimental project with potential but no demonstrated commercial value. It lacks the data needed to assess investment or partnership viability. Further due diligence would require evidence of early adoption, user feedback, or product development milestones beyond the hackathon submission.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.