OpenAI 2026 hackathon

Slide-of-Life

He's after me! The terminator! Quick! Remember: we can go beyond the filenames and file hashes. If he succeeds, then pathologists will lose the one thing saving them from a headache: a Slide-of-Life.

Solo project by Prasiddh Pandya · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #6,769 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Slide-of-Life is a Python command-line tool developed by one person (Prasiddh Pandya) that audits computational-pathology datasets for hidden relationships between training and test sets. It checks for duplicate files, identical content, similar images, and overlapping patient/specimen identifiers across datasets.

What changed

The project was submitted to the OpenAI 2026 hackathon as a self-contained tool built using Python, Codex, and GPT-5.6. It is described as a proof-of-concept for detecting data leakage in AI model training pipelines.

Single most important open question

Is there any evidence of real-world usage or adoption by researchers or institutions? The description states no revenue, customers, or traction data are available beyond the author’s own account.

Back to contents

What The Product Actually Is

The description states that Slide-of-Life is a Python command-line tool designed to audit computational-pathology datasets. It reads spreadsheet manifests (dataset manifests) and identifies connections between training and test sets based on shared identifiers like patients, specimens, or microscope slides.

It detects:

  • Files with identical content
  • Same image saved in different formats
  • Images that look unusually similar

The tool evaluates findings against user-defined rules called a split policy to determine if relationships are permitted or require human review. It generates structured output files and spreadsheets for further analysis, but never modifies the original data automatically.

It was built using:

  • Python
  • Codex (for engineering tasks)
  • GPT-5.6 (optional assistant for understanding spreadsheet columns)

The tool is intended to be run locally by researchers on their own datasets.

Inference The product appears to be a data integrity audit tool aimed at preventing accidental data leakage in AI model development, particularly within pathology research contexts.

Back to contents

Positioning & Claim Evolution

The description states that the inspiration came from exploring pathological datasets and noticing repeated files used for training and testing AI models. This led to the idea of building a tool that goes beyond conventional duplicate detection methods.

It claims to:

  • Analyze computational-pathology datasets
  • Detect hidden relationships between training and test sets
  • Flag potential data leakage issues without modifying original data
  • Provide structured reports and repair suggestions

There is no indication of broader positioning beyond this niche use case in pathology research. The author does not claim it works across industries or platforms, nor does it appear to be marketed as a commercial product.

Inference Slide-of-Life positions itself as an internal tool for researchers working with AI datasets in computational pathology, focused on preventing data leakage rather than general-purpose file management or analysis.

Back to contents

Target Customer & ICP

The description states that the tool is intended for researchers working with computational-pathology datasets, particularly those involved in training and testing AI models. It is designed to help detect when training and test sets contain overlapping patient or specimen data, which can lead to misleading performance metrics.

It is described as a command-line tool meant to be run locally by users on their own datasets.

There is no evidence of:

  • Specific customer segments beyond researchers
  • Industry verticals beyond pathology
  • Any targeting of commercial entities or enterprises

Inference The ICP appears to be individual researchers or small teams in computational pathology who need to validate dataset integrity before model training.

Back to contents

Business Model & Pricing Evidence

The description does not mention any pricing, licensing, or monetization strategy. It is described as a command-line tool built for personal use, with no indication of paid features or subscriptions.

It was published to PyPI (Python Package Index), suggesting it may be freely available for installation via pip.

There is no evidence of:

  • Revenue streams
  • Paid versions
  • Subscription models
  • Customer contracts

Inference No business model or pricing information is evident. The tool seems to be open-source or free-to-use, likely distributed through PyPI.

Back to contents

Technical & Delivery Signals

The tool was built as a Python command-line application, using:

  • Codex for development assistance
  • GPT-5.6 (optional) for column interpretation
  • Standard Python libraries including pydantic, pytest, jinja, pillow, etc.
  • Schema mapping to standardize data formats
  • Perceptual hashing and SHA-256 checksums for file comparison

It reads spreadsheets, cleans identifiers, and compares them against internal logic. It outputs structured files and reports in formats readable by browsers or other software.

The author mentions challenges with packaging and publishing to PyPI, which were resolved using Codex.

Inference The tool is technically feasible and built with standard Python practices, though it has not yet been validated in production environments or widely adopted.

Back to contents

Traction & Maturity Signals

There is no evidence of:

  • Customers
  • Revenue
  • User adoption
  • Product usage metrics
  • Public deployment or integration into workflows

The project was submitted to a hackathon and published to PyPI, but there is no indication that it has been used in real-world research settings.

It is described as a proof-of-concept tool with future plans for testing on larger datasets and improving usability.

Inference The tool exists only as a prototype or early-stage product. No traction or maturity indicators are evident beyond its creation and submission to a hackathon.

Back to contents

Competitive Context

The description does not provide any information about:

  • Competitors
  • Existing tools in the same domain
  • Market landscape
  • Prior art or similar solutions

It is unclear whether there are other tools for detecting data leakage in AI datasets, especially within computational pathology.

Inference No competitive context is provided. The tool may be unique or novel in its approach, but this cannot be confirmed without external references.

Back to contents

Key Risks & Red Flags

  • No real-world usage: The tool has not been tested or used beyond the author’s own environment.
  • Single-person development: Only one team member (Prasiddh Pandya) is involved, which raises concerns about scalability and long-term maintenance.
  • Limited scope: It is only described as useful for computational pathology datasets; no indication of cross-domain applicability.
  • Unverified claims: The tool’s effectiveness in detecting data leakage is not demonstrated or validated.
  • Dependency on GPT-5.6: While optional, reliance on a proprietary AI assistant may raise concerns about reproducibility and consistency.

Inference Risks include lack of validation, limited functionality, and no evidence of commercial viability or adoption.

Back to contents

Diligence Questions To Ask The Founders

  1. Has the tool been tested on actual datasets used in research settings?
  2. What specific types of data leakage does it detect, and how accurate is its detection?
  3. Are there any known limitations or edge cases where it fails to detect issues?
  4. How do you plan to scale beyond a single developer?
  5. Have you considered integrating with existing AI development platforms or pipelines?
  6. Is there any feedback from researchers who have tried the tool?
  7. What are your plans for ongoing maintenance and updates?

Back to contents

Investment/Partnership Verdict

Not evidenced: There is no evidence of revenue, customers, traction, or commercial viability beyond the author’s own description.

The project appears to be a proof-of-concept tool developed during a hackathon, with no indication of real-world application or market demand. It is described as a command-line utility for computational pathology researchers and lacks any signs of product-market fit or monetization strategy.

Confidence level: Low — based entirely on self-reported information without external validation or usage data.

Verdict: Not ready for investment or partnership at this stage. Further evidence of real-world testing, adoption, or traction is required before considering deeper due diligence.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.