OpenAI 2026 hackathon

Indian Law 100: Legal AI Benchmark

An auditable benchmark that tests AI models across 100 Indian-law micro-cases, scoring outcomes, legal authority, reasoning, and resistance to outdated-law traps.

Solo project by Rohas Nagpal · 1 likes · 0 comments

Archive position — measured, not model output

1 like on Devpost

506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #1,223 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

The company appears to be a solo project (1 person) that builds an AI benchmark for evaluating legal reasoning in Indian law. The author states the goal is to test whether AI models can apply current Indian law to realistic legal scenarios rather than simply recall legal knowledge.

Key changes

  • The project was submitted to the OpenAI 2026 hackathon.
  • It focuses on a specific legal domain (Indian law) and uses a structured micro-case approach to evaluate AI performance.
  • It includes technical features like API integration, leaderboard, and checkpointing for long benchmark runs.

The single most important open question

Is there any evidence that this benchmark is being used or adopted by others beyond the author's own testing? The description does not indicate any traction or external usage.

This analysis is based entirely on the self-reported project description provided. No independent verification, revenue data, customer base or adoption metrics are available.

Back to contents

What The Product Actually Is

The description states:

  • Indian Law 100 is a "benchmark that tests AI models across 100 Indian-law micro-cases"
  • It evaluates models using 100 legal micro-cases covering 13 areas of Indian law
  • Each case is scored on three dimensions: Outcome (50%), Legal basis (30%), and Reasoning (20%)
  • The benchmark includes 40 easy, 40 medium, and 20 hard cases
  • Cases deliberately contain legal traps involving repealed statutes, procedural changes, amendments, limitation periods, and outdated precedents
  • It connects to OpenRouter for testing various AI models
  • Results are stored server-side with authenticated access and a public leaderboard
  • Features include automatic progress recovery using IndexedDB, JSON export of sessions, and manual review tools

Back to contents

Positioning & Claim Evolution

The description states:

  • The project was inspired by the need to move beyond legal AI benchmarks that only test recall or exam-style questions
  • It aims to measure whether AI models can "apply current Indian law to realistic legal scenarios rather than simply recall legal knowledge"
  • The author positions it as addressing a specific problem in India following changes to key legal codes (Bharatiya Nyaya Sanhita, Bharatiya Nagarik Suraksha Sanhita, Bharatiya Sakshya Adhiniyam)
  • It is described as an "auditable benchmark" that records every answer, score, judge explanation, and manual adjustment
  • The author claims the dataset and evaluation methodology are "at least as important as the software itself"
  • Future plans include expanding to more legal areas and other jurisdictions

Back to contents

Target Customer & ICP

The description states:

  • The primary users appear to be developers, researchers, and legal professionals who want to evaluate AI systems
  • It is designed for those interested in measuring how well AI applies real-world law—not just how well they answer legal exam questions
  • The benchmark is intended for testing various AI models via API connections (specifically OpenRouter)
  • It targets users who care about transparency and auditability of AI legal reasoning

Not evidenced: Specific customer segments, personas, or use cases beyond the stated audience.

Back to contents

Business Model & Pricing Evidence

The description states:

  • The project connects to OpenRouter for testing AI models
  • Results are stored server-side with authenticated access
  • There is a public leaderboard for comparing model performance
  • No pricing information or monetization strategy is mentioned
  • The author mentions API cost as part of the output but does not specify costs

Not evidenced: Any business model, pricing structure, revenue streams, or commercial arrangements.

Back to contents

Technical & Delivery Signals

The description states:

  • Built with HTML, CSS, JavaScript, PHP, and IndexedDB
  • Connects to OpenRouter for AI model testing
  • Uses REST API architecture
  • Includes features like automatic progress recovery using IndexedDB
  • Supports JSON export of complete benchmark sessions
  • Has authenticated server-side result storage
  • Provides manual review tools and score overrides
  • Implements validation, repair logic, and manual review tools for judge models

Back to contents

Traction & Maturity Signals

The description states:

  • The project was submitted to the OpenAI 2026 hackathon
  • It includes features like checkpointing and recovery for interrupted runs
  • There is a public leaderboard for comparing model performance
  • The author mentions expanding the benchmark with additional legal micro-cases
  • It has been tested with multiple AI models via OpenRouter

Not evidenced: Any usage metrics, user base, revenue, or adoption beyond the author's own testing.

Back to contents

Competitive Context

The description states:

  • Legal AI benchmarks often test recall or exam-style questions
  • This project aims to be different by focusing on applying current Indian law to realistic scenarios
  • It is positioned as addressing a gap in traditional legal benchmarks that don't measure real-world application
  • The author mentions future expansion to other jurisdictions, suggesting awareness of broader market

Not evidenced: Specific competitors, market size, or competitive positioning beyond the stated differentiation.

Back to contents

Key Risks & Red Flags

The description states:

  • The project is built by a single person (Rohas Nagpal)
  • It was submitted as a hackathon entry
  • No evidence of commercial traction or adoption
  • The author acknowledges challenges in designing good benchmark questions and AI-based evaluation
  • The project relies on OpenRouter for model testing, which may limit control over the platform

Key risks:

  • Single-person development limits scalability and support
  • Lack of external validation or adoption suggests limited market demand
  • Reliance on OpenRouter for API access creates dependency risk
  • No clear path to monetization or commercial viability

Back to contents

Diligence Questions To Ask The Founders

  1. What specific legal domains beyond Indian law are you planning to expand into?
  2. How do you plan to validate the accuracy of your benchmark cases and scoring methodology?
  3. Have you received any feedback from developers, researchers, or legal professionals who have used the benchmark?
  4. What is your strategy for scaling beyond a single-person development effort?
  5. How do you plan to monetize this tool if at all?
  6. What are the technical limitations of relying on OpenRouter for model testing?
  7. Have you considered how to handle jurisdictional differences in legal frameworks when expanding internationally?

Back to contents

Investment/Partnership Verdict

The description states:

  • This is a solo project built by one person (Rohas Nagpal)
  • It was submitted as a hackathon entry
  • No evidence of revenue, customers, or traction beyond the author's own testing
  • The author has not indicated any commercialization plans or monetization strategy

Verdict: Not evidenced

This appears to be an experimental project with no demonstrated commercial viability or market traction. The lack of any evidence for adoption, revenue, or customer base makes it difficult to assess investment potential or partnership value. The single-person development and hackathon origin suggest limited scalability or long-term commitment.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.