OpenAI 2026 hackathon

Doclyt: AI Document Restoration

AI-powered document restoration that converts distorted mobile photos into clean, structured documents. Beyond OCR, Doclyt faithfully preserves layout, tables, logos, signatures, and stamps.

Solo project by Arpan Mondal · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,772 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be: Doclyt is an AI-powered document restoration tool that converts distorted mobile photos into clean, structured digital documents while preserving layout, tables, logos, signatures, and stamps. It is described as a system that treats documents as more than text, focusing on visual fidelity and structural accuracy over raw OCR extraction.

What changed: The project was submitted to the OpenAI 2026 hackathon by one founder, Arpan Mondal. No prior version or evolution is described; this is a self-contained, early-stage prototype built for a hackathon context.

Single most important open question: Is there evidence of traction, revenue, or customer adoption beyond the author's own description?

Back to contents

What The Product Actually Is

The description states that Doclyt transforms imperfect mobile photos into clean, structured digital documents while preserving:

  • Document layout
  • Tables
  • Logos
  • Signatures
  • Stamps
  • Reading order
  • Visual hierarchy

It is described as a system that does not generate missing content but instead preserves what is observed in the source image. The author emphasizes that it is designed to be "faithful" rather than "hallucinatory."

The system uses classical computer vision and AI-powered document understanding, including:

  • Image preprocessing (perspective correction, illumination normalization, denoising)
  • OCR and layout analysis
  • Structural understanding of the document
  • Appearance-aware reconstruction
  • Asset preservation for signatures, stamps, and logos
  • Faithful rendering into a clean digital document

Built with FastAPI, Next.js, OpenCV, PaddleOCR, Python.

Inference: The product is a document restoration engine focused on visual fidelity and layout preservation. It is not a general OCR tool but a specialized system for reconstructing documents from low-quality images.

Back to contents

Positioning & Claim Evolution

The author states that Doclyt is not an OCR solution in the traditional sense, but rather a system that treats a document as more than text. The positioning centers on:

  • Faithful preservation of original appearance
  • Visual identity (logos, stamps, typography)
  • Layout and structure fidelity

It is described as being built to "faithfully reconstruct the document—not simply perform OCR."

Inference: The product positions itself as a niche solution for users who need high-fidelity document restoration, not just text extraction. It implies a shift from traditional OCR toward visual and structural accuracy.

Back to contents

Target Customer & ICP

The description does not name specific customers or personas. However, the author states:

  • "We're continuing to improve visual fidelity by refining typography, spacing, tables, and asset restoration while keeping the system deterministic and trustworthy."

This suggests a potential target audience of:

  • Businesses
  • Governments
  • Archives
  • Individuals

Who need to digitize documents with higher fidelity than traditional OCR workflows.

Inference: The ICP is likely users who prioritize visual accuracy over raw text extraction, such as archivists, legal professionals, or government agencies handling official records.

Back to contents

Business Model & Pricing Evidence

No evidence of pricing, monetization strategy, or business model is provided in the description. The project is presented as a hackathon submission with no indication of commercial viability or revenue streams.

Not evidenced

Back to contents

Technical & Delivery Signals

The system is built using:

  • FastAPI
  • Next.js
  • OpenCV
  • PaddleOCR
  • Python

It uses a pipeline that includes image preprocessing, OCR and layout analysis, structural understanding, appearance-aware reconstruction, and asset preservation.

The author notes that the hardest challenge was preserving document identity rather than reading text, and that they spent time balancing enhancement with preservation.

Inference: The technical stack suggests a backend API with frontend UI, using Python-based tools for image processing and OCR. The pipeline is described as deterministic and faithful to input data.

Back to contents

Traction & Maturity Signals

The project was submitted to the OpenAI 2026 hackathon by one founder, Arpan Mondal. No evidence of traction, customers, or adoption beyond the author's own description is provided.

Not evidenced

Back to contents

Competitive Context

No mention of competitors or market positioning is included in the description. The author does not reference existing tools or platforms that offer similar functionality.

Not evidenced

Back to contents

Key Risks & Red Flags

  • Early-stage prototype: Submitted to a hackathon, with no evidence of prior versions or commercial development.
  • No traction or revenue: No data on users, customers, or monetization.
  • Single founder: The team size is listed as 1.
  • Unproven market demand: No evidence that the target audience has a demonstrated need for this specific solution.
  • Self-reported only: All claims are unverified and based on the author’s own description.

Inference: This appears to be an experimental or exploratory project with no commercial validation. It lacks any signal of product-market fit or scalability.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific use cases have you identified for this tool, and how do they differ from existing OCR tools?
  2. Have you tested the system on real-world documents beyond the scope of a hackathon?
  3. Is there any internal data or feedback showing that users value visual fidelity over raw text extraction?
  4. How do you plan to scale this solution beyond a prototype?
  5. What is your roadmap for monetization, if any?

Back to contents

Investment/Partnership Verdict

The description presents Doclyt as an early-stage hackathon project with no evidence of traction, revenue, or customer adoption. It is self-reported and unverified.

Not evidenced

Confidence: Low. The project is described as a prototype built for a hackathon by one person. There is no indication of commercial viability, market demand, or product development beyond the initial submission.

Inference: This is not a viable investment or partnership opportunity at this stage. It may be a promising idea with potential, but lacks any evidence of execution or traction to support further diligence.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.