Varun Goyal · San Francisco

I build the systems that make AI useful, measurable, and better over time.

I’m a founding engineer at WisdomAI. I work around the AI analyst: harnesses that plan and execute, evaluations that explain failure, and feedback loops that turn real use into better systems.

Right now, I’m focused on what happens after an AI analyst answers. Why did it choose that plan? Which piece of context mattered? Where did confidence outrun correctness? I turn those traces and user reports into changes we can test across the system.

Most failures cross boundaries between models, tools, data, and product assumptions. The work is part debugging, part measurement, and part sitting with the people who use the product until the real problem stops hiding.

Orchestration is product.

Planning, tool use, recovery, attribution, and the answer contract determine whether an AI system is useful. The harness is not glue around the model; it is much of the product.

Aggregate scores hide the work.

A useful evaluation exposes the failure modes beneath the average, traces them to system behavior, and makes the next engineering decision clearer.

Context is local.

A company’s data carries its own definitions, exceptions, permissions, and history. Generic schemas rarely capture enough of what a table, column, or document actually means.

This section is intentionally incomplete. More details will appear as I find honest ways to describe the work without publishing the work.

San Francisco

WisdomAI

Founding engineer. Building the harness, evaluation, and improvement systems behind an AI data analyst, with earlier work across unstructured analytics and extraction.

Solana Labs

Solana, for fun

Jumped into the deep end of crypto and helped build Solana’s ChatGPT plugin for wallets, tokens, NFTs, and transactions. A short, slightly chaotic experiment in making blockchains conversational.

IIT Kanpur

Predicting deforestation from space

With Google AI funding, I turned roughly a terabyte of Google Earth and Landsat history into a deforestation forecasting pipeline on machines much smaller than the data. Built with IIT Kanpur, moja global, and Vector AI, the CNN-LSTM model was later presented to government stakeholders in Uganda.

Technical report ↗
During COVIDRubrik

Permissions at a cybersecurity company

At Rubrik, before its IPO, I asked for more than the usual intern-sized project and delivered an overhaul of permissions infrastructure for a subset of cloud backups. COVID definitely gave me the extra time I needed to handcraft secure code back then.

Before that

A few different kinds of systems

Quantitative researcher at Quadeye and grounded language-model applications at Leena AI.

2022–24

MS, Computer Science
University of Illinois Urbana-Champaign

2018–22

B.Tech, Computer Science
IIT Kanpur

2023–24

ICPC & student groups
Represented UIUC at the ICPC Mid-Central Regional; previously helped run IIT Kanpur’s Programming Club and Consulting Group.

Olympiads

Gold medalist, Physics and Junior Science Olympiad Camps.

Reliable systems are more interesting than impressive demos.

A company’s data has local meaning. Generic schemas rarely capture enough of it.

Sitting with the person using the thing is usually faster than guessing what they need.

I’m interested in people building ambitious AI systems where technical depth, product judgment, and customer reality all matter, especially around agent reliability, evaluation, context, and long-lived systems.

Compare notes on LinkedIn