OpenAI 2026 hackathon

Family

Family is a continuous-improvement layer for Codex that learns how humans and Codex can work better together. Family brings rigorous evals to individuals by leveraging their unique tasks.

Solo project by Zane Peycke · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #4,047 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Family is a self-reported private continuous improvement layer for Codex, built as a plugin that enables humans and AI (specifically Codex and OpenAI models) to collaborate more effectively. It uses reinforcement learning (RL) environments to evaluate and improve human-AI collaboration through personalized, local, and privacy-safe evaluation workflows.

What changed

The project is described as an experimental tool developed for the OpenAI 2026 hackathon. It introduces a novel approach to personalizing AI interaction by creating private evals that evolve with user behavior, using deterministic scans and local processing. The author states it was built in Codex using GPT-5.6 over several projects and threads.

Single most important open question

Is there evidence of actual usage or adoption beyond the hackathon context? The description does not indicate any real-world deployment, customer base, or revenue — only a self-reported prototype with early tester benchmarks.

Back to contents

What The Product Actually Is

The description states that Family is:

  • A private continuous improvement layer for Codex.
  • A personalized private eval in the form of an RL environment.
  • A Codex plugin, built using gpt-5.6, Python 3.12+, and uv workflow engine.
  • Designed to allow users to measure how factors like reasoning effort, skills, and goal setting impact their work.
  • Capable of turning representative work into a reviewed RL environment using PrimeIntellect.
  • A tool that supports both:
    • Users already using Codex: by enabling local activity review within Codex.
    • New users to Codex: by offering one bounded first task, practice, and visible evidence.

Inference Family appears to be a developer-focused tool aimed at improving human-AI collaboration through structured evaluation and feedback loops. It is not described as a general-purpose AI assistant or marketplace but rather a personalization layer for AI tools like Codex.

Back to contents

Positioning & Claim Evolution

The author states:

  • Family aims to bring rigorous evals to individuals by leveraging their unique tasks.
  • It seeks to lay the foundation for increasingly capable AI systems with greater personalization and safety.
  • It is positioned as a continuous improvement layer, not just a static tool.
  • The product supports both experienced Codex users and newcomers.

Inference The positioning evolves from a hackathon prototype to a vision of personalized, safe, and scalable AI interaction. However, the claims are largely aspirational — there is no evidence of real-world traction or validation beyond early testers.

Back to contents

Target Customer & ICP

The description states:

  • Target users include:
    • Those already using Codex.
    • New users to Codex who want a guided start path.
  • It supports two user types:
    • Users with existing Codex history (for review and improvement).
    • Users without prior history (who are given one task, practice, and evidence).

Inference The ICP appears to be developer tooling users, particularly those working in AI-assisted environments like Codex. The product is not described as targeting end-users or non-technical teams.

Back to contents

Business Model & Pricing Evidence

Not evidenced.

The description does not mention:

  • Any pricing model.
  • Revenue streams.
  • Monetization strategy.
  • Subscription plans or usage-based billing.

Inference There is no indication that Family has a business model beyond its hackathon prototype. It is described as a plugin, but no commercial structure is implied.

Back to contents

Technical & Delivery Signals

The description states:

  • Built in Codex, using gpt-5.6, Python 3.12+, and uv.
  • Implements:
    • Consent-first local evidence inventory.
    • Privacy-safe summaries.
    • Representative authoring planner (stores metadata, not raw prompts).
    • Trusted synthetic compiler families: repository change monitoring and bounded evidence briefs.
    • Native verifiers (0.2.0 Taskset, Codex Harness artifacts).
    • Actor/host package separation, mutation proofs, exact regeneration, canary tripwires.
    • Reversible improvement previews, approvals, task bindings, results, rollback receipts.
    • Reproducible stripped private-preview bundles with checksums and activation controls.
    • CI across Python 3.12 and 3.13 covering tests, linting, packaging, plugin validation, offline demo, and sanitized bundle.

Inference The technical architecture is described as robust, privacy-focused, and built for deterministic workflows. It includes strong safety and reproducibility features, but these are not validated in real-world usage.

Back to contents

Traction & Maturity Signals

Not evidenced.

The description does not include:

  • Any customer data.
  • Revenue figures.
  • Adoption metrics.
  • Product usage statistics.
  • Real-world deployment or pilot programs beyond the hackathon.

Inference There is no evidence of traction or maturity beyond an experimental prototype. The early tester benchmarks are mentioned, but no data on scale or impact is provided.

Back to contents

Competitive Context

Not evidenced.

The description does not:

  • Name competitors.
  • Describe market positioning relative to other AI collaboration tools.
  • Mention any competitive advantages or differentiators.

Inference No competitive context is provided. The product is described in isolation, with no reference to existing tools or markets.

Back to contents

Key Risks & Red Flags

The description states:

  • The hardest problem was deciding when evidence deserved to count, especially around privacy and evaluation stability.
  • It cannot treat missing history as a deficiency — the beginner path starts from one desired outcome.
  • Evaluation stability belongs in the product experience, not in automated metrics.

Inference

  • Risk of limited adoption due to complexity or lack of clear value for new users.
  • Risk of low scalability if personalization is too manual or requires deep user involvement.
  • Red flag: The project is described as a hackathon prototype with no commercial traction or validation.

Back to contents

Diligence Questions To Ask The Founders

  1. What are the actual early tester benchmarks, and how were they collected?
  2. How does Family handle privacy in real-world usage — especially when users have sensitive prompts or code?
  3. Is there any plan to expand beyond Codex or integrate with other AI tools?
  4. What is the roadmap for moving from a hackathon prototype to a product with real users?
  5. Are there any commercial partnerships or pilot programs already underway?

Back to contents

Investment/Partnership Verdict

Not evidenced.

The description does not provide:

  • Any financial data.
  • Evidence of traction or revenue.
  • Customer validation.
  • Market opportunity or competitive positioning.

Inference This is a pre-product, pre-traction prototype, submitted for a hackathon. There is no evidence to support an investment or partnership decision at this stage. The project shows technical ambition and a clear vision but lacks commercial proof of concept.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.