OpenAI 2026 hackathon

Penny

Penny is an Azerbaijani AI finance assistant that turns everyday text and voice into structured expenses and clear reports, backed by a safe eval lab that tests every interpretation.

Solo project by ElmaddinSuleyman Suleyman · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #5,884 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Penny is an Azerbaijani-language AI finance assistant that allows users to record expenses using natural language (text, voice, receipts), with a focus on simplicity and safety. The project includes a standalone evaluation lab for testing AI interpretations before they reach real users.

What changed

During Build Week, the team added a "Penny Trust Preview & Eval Lab" — an isolated system that evaluates how GPT-5.6 Sol interprets synthetic Azerbaijani financial messages without executing any actions. This lab tests intent, amount, category, merchant, date, correction requests, and ambiguity, but does not write or modify real data.

Single most important open question

Is there evidence of traction, revenue, or user adoption beyond the author’s own development work?

Back to contents

What The Product Actually Is

The description states that Penny is an Azerbaijani-language AI finance assistant. It supports:

  • Recording expenses via text input
  • Submitting voice notes
  • Extracting expense data from receipt images
  • Correcting previously recorded expenses
  • Searching transactions
  • Generating clear spending reports

It operates through a Telegram interface.

Additionally, during Build Week, the team introduced a standalone evaluation lab that tests AI interpretations in isolation. This lab:

  • Evaluates synthetic Azerbaijani financial messages using GPT-5.6 Sol
  • Displays and validates fields such as user intent, amount, currency, category, merchant, date, correction targets, ambiguity, and whether the message represents a write action
  • Does not execute any changes; all results are non-executable candidates
  • Is completely separated from production systems (database, Telegram flow)

The lab uses strict structured outputs, deterministic validation, and privacy controls to ensure safety and reproducibility.

Claim

The product is an AI finance assistant with a focus on natural language interaction and financial safety.

Fact

The description states this. No evidence of actual users or usage exists beyond the author’s own development work.

Back to contents

Positioning & Claim Evolution

The project positions itself as a natural-language expense-tracking tool for Azerbaijani speakers, aiming to simplify financial record-keeping by removing form-filling steps.

It emphasizes:

  • Simplicity: Users interact through conversation-like inputs
  • Safety: AI interpretations are tested in an isolated lab before reaching real users
  • Localization: Specifically built for Azerbaijani language and context

The evolution from basic expense tracking to including a Trust Preview & Eval Lab shows a shift toward building trust and safety into the AI interpretation process.

Claim

Penny aims to be a safer, more intuitive financial assistant.

Fact

The description states this. No evidence of market positioning or user feedback is provided.

Back to contents

Target Customer & ICP

The target customer appears to be Azerbaijani speakers who need to track personal expenses using natural language methods.

The product is designed for users who prefer conversational interaction over traditional form-based expense tracking.

There is no indication of B2B use cases or enterprise adoption.

Claim

The product targets Azerbaijani individuals seeking simplified expense management.

Fact

The description states this. No evidence of actual customers or user segments is provided.

Back to contents

Business Model & Pricing Evidence

No information about pricing, monetization, or business model is included in the description.

The project is described as being in a close beta phase, and no revenue data, customer acquisition costs, or monetization strategy are mentioned.

Claim

The business model is unclear.

Fact

Not evidenced. The description does not state any commercial details.

Back to contents

Technical & Delivery Signals

Key technical elements include:

  • Use of GPT-5.6 Sol for interpretation
  • Codex for engineering and testing
  • Python-based implementation
  • Telegram interface
  • Strict structured outputs via OpenAI API
  • Offline, isolated evaluation lab
  • Deterministic validation and schema enforcement
  • Synthetic fixture generation
  • GitHub Actions for CI/CD

The system is built to prevent silent fallbacks, protect privacy, and avoid production contamination.

Claim

The product uses advanced AI with strong engineering safeguards.

Fact

The description states this. No evidence of deployment, scalability, or performance metrics is provided.

Back to contents

Traction & Maturity Signals

The project is described as being in a close beta phase.

It includes:

  • 17 synthetic Azerbaijani financial fixtures
  • 36/36 focused tests passed
  • 17/17 offline evaluation cases passed
  • 53/53 standalone test suite passed
  • One controlled live sample of GPT-5.6 Sol integration

However, there is no evidence of:

  • Real users or customer base
  • Revenue or monetization
  • Product adoption or usage metrics
  • Production deployment or scaling

Claim

The project has demonstrated technical capability.

Fact

The description states this. No evidence of traction or real-world impact is provided.

Back to contents

Competitive Context

The description does not mention any competitors or market context.

It focuses on the unique aspects of the product (natural language, safety lab) but does not discuss how it compares to existing expense-tracking tools or AI finance assistants.

Claim

The competitive landscape is unknown.

Fact

Not evidenced. No comparison or market positioning data is provided.

Back to contents

Key Risks & Red Flags

  • No traction or user adoption: The product is described as in close beta with no evidence of real users.
  • Unproven commercial viability: No pricing, revenue, or monetization strategy is evident.
  • Limited scope: Only one developer is involved; no team expansion or external validation.
  • Self-contained development: The evaluation lab is isolated and not integrated into production.
  • No third-party verification: All claims are self-reported with no independent confirmation.

Inference Without real users, revenue, or market traction, the project may be in early-stage development without commercial viability.

Fact

Not evidenced. These are risks inferred from lack of evidence.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the plan for moving from close beta to a public product?
  2. How will you validate that the AI model correctly interprets real-world inputs, not just synthetic ones?
  3. Are there any plans for monetization or revenue generation?
  4. What are the next steps in scaling beyond one developer and one evaluation lab?
  5. How do you intend to onboard users and gather feedback from real-world usage?
  6. Is there any evidence of user interest or demand outside of the hackathon context?

Back to contents

Investment/Partnership Verdict

At this stage, Penny is a self-contained prototype built during a hackathon with strong engineering rigor in its evaluation lab.

It shows potential for building a safer AI-driven expense assistant, but lacks:

  • Real users
  • Revenue or monetization
  • Market traction
  • Clear commercial strategy

Verdict Not ready for investment or partnership without further evidence of traction, user adoption, or business model development.

Confidence Level Low — based on self-reported evidence only.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.