AI Robot Operator at Physical Intelligence
MS in Software Engineering · BS in Computer Science
I build systems across robotics, machine learning, AI reliability, and full-stack software. My recent work focuses on robot-fault diagnostics, demonstration-data quality, reproducible ML evaluation, and scientific benchmark auditing.
-
Curation Metrics Are a Post-Training Decision: Auditing Trajectory Smoothness for Adapting Robot Foundation Models: Accepted as a poster at the NeurIPS 2026 RoboPAD workshop. Builds on the base preprint Smoothness Ranks Skill, Not Success: An Audit of Published Trajectory-Smoothness Curation Metrics Under Episode-Length Controls; the CoRL WEBP and CoRL Oops, I Erred versions remain under review.
-
Not All Bad Demonstrations Are Equally Bad: Quantifying How Demonstration Failure Modes Degrade Closed-Loop Policy Performance: Controlled study of 660 policy fits across five failure modes, two policy families, ten seeds, and transition-matched controls. Under review at the CoRL 2026 Learning from Corrections and Interventions workshop (OpenReview).
-
What Can Three Seeds and Fifty Rollouts Resolve? A Preregistered Audit of Operator and Seed Variance on robomimic Multi-Human Benchmarks: Preregistered audit of 138 BC-RNN runs splitting robomimic Multi-Human benchmark variance into seed, demonstration, operator, and rollout noise. At 50 rollouts, ten seeds were statistically indistinguishable on every task; re-evaluating 46 Square checkpoints at 500 rollouts made the seed effect detectable and erased an apparent operator-proficiency pattern. Under review at the CoRL 2026 SimBench2Real workshop (OpenReview), with a preprint and Zenodo archive.
-
One Byte, One Rank Reversal: Auditing HumanEval Pipeline Sensitivity for Base Code Models: A pinned five-model reproduction on an RTX 4070 found StarCoder2-3B scoring 3/164 instead of its published 31.7%, reversing a published ranking. Removing one trailing newline from the evaluation prompt raised it to 49/164 and restored the order. Technical report v1.9 with a versioned Zenodo archive. PDF
-
When Does INT8 Actually Accelerate Host-CPU Inference? A Reproducible Audit of MobileNetV2 Conversion and Packaging: Shows that INT8 latency depends on backend delegation and quantization granularity, alongside paired accuracy and packaging controls. Released as v1.2.0 on Zenodo, not yet submitted for peer review.
-
Conditioning and Directionality in Adversarial Transfer on CIFAR-10: Denominator-aware full-test study across three architectures and three training seeds. Under review at Transactions on Machine Learning Research (OpenReview), with an archived reproducibility release.
-
Complete Patient-Level Leakage Among Traceable Images in a Widely Used Brain Tumor MRI Benchmark: Finds complete patient overlap among traceable tumor test images plus substantial duplicate contamination, and releases patient-disjoint folds. Available as a preprint and Zenodo artifact, not yet submitted to a journal.
-
Sampling Budget Biases Apparent Metastable-State Counts in Molecular Simulation: Controlled synthetic landscapes and two independent alanine-dipeptide trajectories show how sampling budget can inflate apparent state counts. Under review at JCTC (submitted September 15, 2026), with a ChemRxiv preprint and a versioned Zenodo dataset.
-
π0 on a Budget — Fine-tuning π0-FAST on a consumer RTX 4070 using synchronized demonstrations from a custom four-degree-of-freedom teleoperation arm, followed by open-loop and closed-loop evaluation.
-
Demonstration Quality Robustness — Controlled simulation study of 660 policy fits measuring how specific demonstration failure modes degrade closed-loop policy performance, and how poorly open-loop evaluation tracks that damage. Paper (PDF).
-
UR5e RTDE Fault-Injection Harness — Drives a simulated UR5e over RTDE, the same interface used by production UR5e cells, logs synchronized joint telemetry, and deliberately induces connection faults to characterize whether the client fails with a clean exception, a hang, or a native crash.
-
openpi Language-Steerability Demo — Runs the public π0.5 LIBERO checkpoint in a single simulated scene with different language instructions and scores success with the simulator's own goal checks. Trained and paraphrased prompts succeed 10/10, while novel object recombinations drop as low as 1/10.
-
Grasp Annotation Dataset Pipeline — Data-ops pipeline that samples public robot-arm video into frames, provisions a Labelbox project with a grasp-focused ontology, and exports completed labels as COCO JSON. Each frame can carry a
grasp_eventbounding box, anobject_contactkeypoint, and afailure_modelabel (drop, miss, collision, timeout, or success). -
CAN Bus Fault Injector — Three-node Arduino/MCP2515 CAN bus with SocketCAN monitoring and a documented catalog of physical and protocol-level faults. Article on Medium.
-
GELLO-Style Leader–Follower Arm — Approximately $30 potentiometer-based teleoperation rig using direct joint-space control without inverse kinematics. Article on Medium.
-
Scribbot — A low-cost, mostly 3D-printed desktop robot arm that draws on a 3 × 3 inch sticky note: an MG90S base on a 2:1 printed gear, MG996R shoulder and elbow, an MG90S wrist, and a gravity pen holder, driven by an Arduino Nano and PCA9685 servo driver, with OpenSCAD sources, printable STLs, and stage-by-stage build guides.
-
Chess-RL Engine — AlphaZero-inspired engine that learns only through self-play: a C++ bitboard move generator exposed through pybind11, a PyTorch policy/value network on an 18-channel board encoding, and a PyGame interface showing move probabilities in real time.
-
Neural Net Mapper — Trains an MLP on a synthetic shapes dataset and renders an animated map of its inner workings: neuron activations, weight signs and magnitudes, dropout, predictions, and live loss/accuracy curves.
-
MedStract — Flask web app that retrieves peer-reviewed abstracts from PubMed and turns them into summaries calibrated to four reading levels, from the general public to the domain expert, with question answering over the results, publication-trend charts, and citation exports (APA, MLA, Chicago, Vancouver, BibTeX, RIS).
-
Castline Studio — Production e-commerce order system with client-side STL analysis, Etsy and shipping integrations, and automated email.
Stanford, CA · September 2026–Present
San Francisco, CA · May 2026–Present
- Generate high-quality demonstrations for general-purpose robotics foundation models by teleoperating robotic arms through manipulation, sorting, and assembly tasks.
- Evaluate model behavior against quality benchmarks, identify edge cases and failure modes, and annotate task data for model development.
- Diagnose recurring hardware–software faults across distributed robotics systems by tracing failures through service logs and documenting root causes for engineering teams.
Remote · April 2026–June 2026
- Evaluated LLM-generated code and technical responses across AI/ML and software-engineering domains, scoring correctness, instruction following, and writing quality.
Berkeley, CA · October 2023–June 2024
- Developed a React, Express, REST, and SQL application for uploading, processing, and reviewing clinical-laboratory instrument data.
- Built validated data pipelines that consolidated more than 50 instrument fields into an automated upload workflow.
- Investigated data-recording defects, documented failure patterns, and improved pipeline reliability.
- Accelerated Computing in Modern CUDA C++ (September 2026)
- HackerRank Software Engineer Certificate (March 2026)
- Cisco Network Support and Security (September 2025)
- IBM Artificial Intelligence Fundamentals (August 2025)
- Stanford Fundamentals of AI and Machine Learning in Healthcare (June 2025)
- Google Foundations of Project Management (September 2024)
- Google IT Support Professional Certificate (May 2021)
- MS, Software Engineering — Grand Canyon University (May 2024–October 2025)
- BS, Computer Science — University of Silicon Valley (May 2021–August 2023)
Open to full-time opportunities in robotics software, AI/ML engineering, full-stack development, and AI reliability.

