OpenAI 2026 hackathon

Batchwork AI

One typed Python API and CLI for provider-native AI batch jobs across OpenAI, Anthropic, Gemini, Groq, Mistral, Together, and xAI saving up to 50% off inference.

Solo project by Ajan Raj · 1 likes · 0 comments

Archive position — measured, not model output

1 like on Devpost

506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #676 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Batchwork AI is a self-reported Python library, CLI, and agent skill that enables users to submit batch jobs across multiple AI providers (OpenAI, Anthropic, Google Gemini, Groq, Mistral, Together, xAI) using a single typed async API. The author states it aims to reduce inference costs by up to 50% compared to synchronous calls, while abstracting provider-specific protocols.

What changed

The project was built as a solo effort over approximately five days, with the author claiming that Codex (a tool for AI-assisted development) was used extensively in its creation. It includes a Python SDK, CLI, and an agent skill for coding agents. The author describes it as a response to personal workflow inefficiencies in using LLMs for classification tasks.

Single most important open question

Is there any evidence of real-world usage or adoption beyond the author’s own development and testing?

Back to contents

What The Product Actually Is

The description states that Batchwork AI consists of:

  • A typed async Python API for submitting text, embedding, or image batches across multiple providers.
  • A CLI tool with commands like submit, run, status, wait, results, cancel.
  • An Agent Skill, which teaches coding agents how to interact with the CLI and includes safeguards such as asking before spending money and resuming interrupted jobs.

The author also mentions that it supports:

  • Normalized job results, usage, and errors correlated by custom_id.
  • Storage, polling, and signed webhooks for production pipelines.
  • Integration with tools like Codex and GPT-5.6 for development.

Evidence All of this is self-reported by the author; no third-party verification or demonstration of product functionality beyond personal use is provided.

Back to contents

Positioning & Claim Evolution

The author positions Batchwork AI as a solution to inefficiencies in using batch APIs across multiple LLM providers. The core claim is that:

  • It reduces inference costs by up to 50%.
  • It simplifies integration work by abstracting provider-specific protocols.
  • It allows users to run batch jobs with minimal effort, e.g., “ask Codex to run it through Batchwork and it costs half as much.”

The author also notes that they were inspired by a TypeScript SDK by Hayden Bleasel but built a Python version with additional features like a CLI and agent skill.

Inference This suggests a niche focus on developers or data scientists who perform classification tasks using LLMs and want to reduce cost and complexity of batch processing.

Back to contents

Target Customer & ICP

The author describes their own use case:

  • They are a data scientist.
  • Their work involves classification tasks, such as labeling support tickets, comments, or transactions.
  • These jobs run offline, without user interaction, and can be costly due to LLM usage.

They state that they personally “barely used” batch APIs because of the integration overhead, which led them to build Batchwork AI.

Inference The target customer appears to be technical users (data scientists, ML engineers) who perform large-scale classification or processing using LLMs and are looking for cost-efficient ways to do so.

Back to contents

Business Model & Pricing Evidence

There is no evidence of pricing, revenue, or monetization strategy in the description. The author does not mention any commercial model, subscriptions, or paid features.

Not evidenced

Back to contents

Technical & Delivery Signals

The author reports:

  • Built with Python and various libraries including asyncio, pydantic, httpx, click, uv, sqlite, etc.
  • Uses Codex and GPT-5.6 for development, including planning, implementation, hardening, verification, and shipping.
  • Includes tests, documentation, and a release pipeline.
  • Supports multiple providers with custom adapters per modality (text, embedding, image).
  • Addresses security concerns such as SSRF protection, credential hygiene, and webhook retries.

Inference The technical stack and delivery process suggest a high degree of automation and tooling integration. However, there is no evidence of production deployment or usage by others beyond the author’s own testing.

Back to contents

Traction & Maturity Signals

There is no evidence of:

  • Customers
  • Revenue
  • Adoption metrics
  • Product usage data
  • Any form of traction beyond personal use and internal testing

Not evidenced

Back to contents

Competitive Context

The author references a TypeScript SDK by Hayden Bleasel, which inspired the Python version. However, there is no mention of other similar tools or platforms in the market.

Inference It seems to be a niche tool addressing a specific gap in batch processing for LLMs across multiple providers. The lack of competitive analysis indicates either limited awareness of existing solutions or no public comparison data.

Back to contents

Key Risks & Red Flags

  • Solo project: Only one team member (the author) is involved, which raises questions about scalability and long-term maintenance.
  • No external validation: No third-party reviews, customer feedback, or usage data are provided.
  • Self-reported only: All claims are unverified; no independent confirmation of performance, cost savings, or real-world utility.
  • Limited scope: The tool is described as solving a specific problem for one person — classification tasks — and may not generalize well to broader use cases.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual cost difference between using Batchwork AI vs. direct provider APIs?
  2. Has the tool been tested in real-world workflows beyond your own?
  3. Are there any known issues or limitations when running large-scale batch jobs?
  4. How do you plan to support additional providers or handle changes in existing provider APIs?
  5. What is your roadmap for product development and maintenance, given that it’s a solo project?

Back to contents

Investment/Partnership Verdict

Not evidenced

The description provides no information on:

  • Revenue
  • Customers
  • Traction
  • Market size
  • Financials
  • Team structure beyond one person

This is a self-reported, unverified, early-stage tool built by a single developer. It has not demonstrated any commercial viability or adoption outside of personal use.

Confidence level Low

Next steps

If this were part of an investment or partnership consideration, further due diligence would require independent validation of the product’s utility, performance claims, and potential for scaling beyond one individual developer.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.