Selected work

Product judgment backed by inspectable evidence

These are not feature inventories. Each case states what I owned, how I reduced uncertainty, what changed, and where confidentiality or study design limits the claim.

Enterprise AI/ML

2022–present

ML-powered spend controls

Product direction for fraud and spend-management systems at Capital One.

My role
Group Product Manager responsible for product strategy, roadmap decisions, and cross-functional alignment across engineering, data science, design, and risk partners.
How I worked
Translate uncertain model behavior and policy constraints into product boundaries, operating decisions, and measurable launch criteria.
What changed
The public evidence is intentionally qualitative: product area, decision ownership, and cross-functional scope are disclosed; confidential customer, risk, and business metrics are not.

0-to-1 production agent

Founded 2025

Honeydew

A family coordination product that turns voice, text, and photo input into calendars, lists, and shared plans.

My role
Founder and accountable product owner across research, strategy, implementation, instrumentation, evaluation, and distribution.
How I worked
Use real operating constraints to define when the agent should act, ask, confirm, or simply answer. Convert recurring failures into synthetic evaluation scenarios before changing behavior.
What changed
A production product that doubles as a feedback system for safer agent decisions. Private prompts, user data, and proprietary workflows remain private.

Agent restraint

June 24, 2026 snapshot

Six models, 227 scenarios

4,086 calls tested how often models act, ask, confirm, or chat when household requests are ambiguous or irreversible.

My role
Defined the product policy, wrote the synthetic scenario set, designed the run and scoring approach, analyzed the results, and published the limitations.
How I worked
Three shuffled trials per model-scenario pair at temperature 0.2, deterministic routing scores, a three-model label panel, and explicit parse-failure accounting.
What changed
Shifted the model decision from a single accuracy ranking to safety, messy-input robustness, steerability, and disagreement on the answer key.

Model selection

April 15, 2026 snapshot

Eight models, 2,800 calls

A bounded production case study comparing model behavior across one prompt and 35 synthetic scenarios.

My role
Framed the selection question, created the benchmark and reporting, and separated reusable evaluation lessons from model-ranking claims the data could not support.
How I worked
Ten trials for every model-scenario pair, with tool-use, parsing, latency, cost, and non-action behavior analyzed separately.
What changed
Found that price and frontier status did not predict fit for this workload, while also documenting prompt coupling and dated provider behavior.

Distribution measurement

January–April 2026 window

What LLM referrals actually showed

A 90-day descriptive field note that kept 13 GA4 sessions and 29 custom events as different units rather than manufacturing a capture rate.

My role
Designed the query audit, instrumented the observable referral path, re-audited the methodology, and narrowed the public conclusion.
How I worked
Separate one ten-query Perplexity panel from ordinary web-search visibility and aggregate referral analytics; publish only descriptive observations.
What changed
Corrected an initially overconfident interpretation and turned the failure into a reusable measurement rule and public correction record.

Publishing system

Open source

PeteWebsite

The canonical portfolio and research archive, built as a small evidence and distribution system rather than a static brochure.

My role
Product owner, writer, maintainer, and accountable editor. AI assists with implementation and review under a public disclosure policy.
How I worked
Typed profile and research registries, automated claim checks, structured data, consent-aware attribution, full-content RSS, and visible correction practices.
What changed
One source of truth now drives public profile copy, machine-readable summaries, research links, and cross-surface consistency.

The through-line

I lead AI/ML products, build production agents, and publish rigorous evaluations of how they behave.