Group Product Manager, AI/ML · Founder · Hands-on builder

Pete Ghiorse

I lead AI/ML products, build production agents, and publish rigorous evaluations of how they behave.

At Capital One, I lead AI/ML product work. Through Honeydew, I build a production family agent and use its failures to sharpen product decisions, evaluation methods, and safeguards.

New York City

Selected evidence

Products, decisions, and inspectable work

See all selected work →

Production product

Built since 2025

Honeydew

A multimodal family agent that turns voice, text, and photos into shared household plans.

My role
Founder and product owner across discovery, product strategy, implementation, evaluation, and distribution.
How I worked
Use production-shaped failures to define decision boundaries, evaluation cases, and safer act/ask/confirm behavior.
What changed
A working product and a repeatable operating loop for improving agent reliability without hiding uncertainty.

Agent evaluation

June 2026 snapshot

When should an agent stop?

227 synthetic household scenarios across 6 models and 4,086 calls.

My role
Designed the product policy, scenario set, measurement approach, analysis, and public correction standard.
How I worked
Repeat each scenario three times; score routing deterministically; retain parse failures, judge disagreement, and limitations.
What changed
Reframed model choice around safety, robustness, and steerability rather than one aggregate accuracy score.

Model selection

April 2026 snapshot

The 2,800-call benchmark

8 models × 35 scenarios × 10 trials, tested against one production-shaped workload.

My role
Turned a model-selection question into a bounded benchmark with explicit tradeoffs and dated conclusions.
How I worked
Compare tool use, restraint, parsing, latency, and cost while separating a dated case study from universal model rankings.
What changed
Exposed prompt coupling and failure modes that price and leaderboard position did not predict.

How I keep the work sharp

Build → Instrument → Evaluate → Publish → Revise

I use shipped systems to expose the next important question, turn failures into evaluation cases, and publish enough method and limitation detail for someone else to inspect the conclusion.

AI product leadership

Strategy, roadmaps, decision boundaries, governance, and cross-functional execution for products that depend on uncertain model behavior.

Agent evaluation & safety

Act/ask/confirm policy, destructive-action guard cases, messy-input robustness, and production-shaped evaluation design.

Multimodal systems

Voice, text, images, tool orchestration, and the product decisions that make an agent useful outside a demo.

Evidence-backed communication

Methods, limitations, negative results, corrections, and visual explanation that let readers inspect the reasoning.

How I use AI and own the final work →

Recent writing

View all

ChatGPeTe

Research editions and the occasional essay.

I lead AI/ML products, build production agents, and publish rigorous evaluations of how they behave—plus essays when I have something worth saying.

Subscribe to ChatGPeTe