Reveal technical notes, experiment results, and ML terminology. Your preference is remembered on this device.

A PERSONAL ML LAB

Saksham
Verma.

Exploring how machines learn.
Building things along the way.

Computer Engineering student at Thapar. Curious about the models underneath, the systems around them, and what happens when you build both.

Explore my work

Ideas, in motion.

Answers you can trace.

Turning financial filings into answers backed by sources.

Explore FinSight

Illustrated workflows

ML / DEEP LEARNING / INTELLIGENT SYSTEMSTHAPAR INSTITUTE · CLASS OF 2027

01 / SELECTED PROJECTS

Ideas, put to work.

A few things I've built.
Each one, a different set of questions.

02 DEEP LEARNING / EDGE AI

SeizureNet.

Smaller models. Real constraints.

Predicting seizure onset from EEG signals, then fitting the model onto a microcontroller. I co-own software and model development in this four-person capstone.

0.88SeizureNet AUROC
395 msTinySeizureNet inference

CHB-MIT · 21 patients · NUCLEO-F446RE deployment

PyTorch1D-CNNEdge deployment
Inside the project

SeizureNet is a 123K-parameter dilated 1D-CNN designed to predict onset 30 minutes ahead. On the CHB-MIT evaluation, it reached 0.88 AUROC, compared with 0.876 for a published 3.2M-parameter reference model.

The evaluation detected 21 of 21 test seizures with a 27-minute average warning and 0.94 false alarms per hour.

TinySeizureNet uses 100K parameters and 54% less peak memory. It runs on a NUCLEO-F446RE microcontroller with bit-identical output and no AUROC loss relative to its software implementation.

Ongoing capstone · October 2025–present

03 AGENTS / MULTI-VIDEO RESEARCH

VidSense.

From watching to asking.

A research agent that searches across videos, compares what they say, and checks answers against evidence. A planner chooses the next tool; a separate verifier checks the claims.

93.8%Agent task success
56.2%Single-shot RAG baseline

One benchmark run · 16 tasks across six videos

LangGraphFAISSFastAPIAzure
Try VidSense (opens in new tab)
Inside the project

The planner can search transcripts, look up timestamps, compare videos, and use web verification. When a tool returns weak evidence or fails, it can retry or switch tools within a six-step budget.

I built a 16-task benchmark across six videos, comparing the agent with single-shot RAG on task success, tool-selection F1, grounded ratio, latency, token use, and estimated cost. The percentages above describe one run of that benchmark.

The FastAPI service streams responses with SSE and is deployed on Azure in Docker.

August 2025–present

02 / EXPERIMENTS

Questions, still open.

Early findings.
Room to change my mind.

EXP. 001 / CODING AGENTS

Paused Initial results available

Coding Agent
Harness Lab.

How much can the harness change a model's coding performance?

I'm exploring the system around a fixed language model: how it receives a problem, uses feedback, and tries again. The first round compares Llama 3.1 8B and Llama 3.3 70B.

Inference access and rate limits interrupted the first round. I plan to revisit model and benchmark choices before continuing.

THE EXPERIMENT

KEEP FIXED
Model weights
CHANGE
Prompts, feedback, retries, context
MEASURE
Correctness, token use, turns
GenerateExecuteFeedback

The loop changes. The model weights stay fixed.

70B / REACT-STYLE PROMPT+4.9 pp

First-attempt success
82.9% → 87.8%

8B / STRUCTURED PROMPT+1.8 pp

First-attempt success
67.1% → 68.9%

8B / CONSTRAINT-ONLY PROMPT−22.0 pp

First-attempt success
67.1% → 45.1%

Preliminary run summaries · 164 tasks per configuration shown · pp = percentage points

EARLY OBSERVATION The same prompt strategy can behave differently across models. More complexity doesn't automatically mean better results.

Read the experiment notes

What this round covers

The first phase was designed around HumanEval+ and its 164 coding tasks. The log contains prompt comparisons and initial feedback/tool experiments. Broader context-management, retry-budget, and automated prompt-search studies remain part of the plan.

How to read the scores

First-attempt success is the log's pass@1. “Within 3 turns” is its reported pass@3; the planned loop uses failure feedback between attempts, so this is presented as iterative success rather than independent-sample pass@k.

First-round results · Run log dated 25 May 2026
ModelConfigurationFirst attemptWithin 3 turnsTokens / task
70BBaseline82.9%93.3%672
70BChain-of-thought84.8%94.5%1,105
70BReAct-style87.8%97.0%970
70BStructured output84.1%94.5%1,046
8BBaseline67.1%77.4%1,169
8BChain-of-thought67.1%78.7%1,827
8BConstraint-only45.1%73.8%1,986
8BReAct-style68.3%76.8%1,523
8BStructured output68.9%77.4%1,524
8BExecution + linter66.5%76.2%1,140
8BExecution + test generation66.5%76.2%1,183
8BVague feedback67.1%73.8%937

Observations worth following

On 8B, chain-of-thought matched baseline first-attempt success while using about 56% more tokens per task. The structured prompt's 1.8-point gain corresponds to roughly three extra solved tasks. These are early observations, not a settled ranking of harness designs.

Limits of this comparison

  • Partial runs (155/164 and 67/164) are excluded.
  • The 70B linter run reported zero tokens and zero success; it is excluded pending investigation.
  • Repeated-run uncertainty is not reported in the log.
  • The two models differ in version as well as size.
NEXT ROUND / NOT STARTED

Revisit model access and benchmark choice, establish fresh baselines, and keep the new results separate from this first round.

03 / ABOUT

Still learning.
Always building.

I'm Saksham, a Computer Engineering student at Thapar Institute of Engineering & Technology, graduating in 2027.

I've been building retrieval systems, AI agents, and models that run on small devices. I'm especially curious about model training: how models learn, how they represent information, and what makes them work.

Away from the code, I follow cricket, tennis, football, and chess. There's usually a match worth keeping an eye on.

B.E. Computer Engineering8.78 / 10 CGPA

What I work with

MODELS & DATA
PyTorch, TensorFlow, scikit-learn, Pandas, NumPy
SYSTEMS & TOOLS
Python, C/C++, SQL, FastAPI, Docker, Redis, Azure
CURRENT INTERESTS
Model training, transformers, efficient deep learning, retrieval, agents

A few milestones

01

Trilemma Beta

Won the global data science tournament, with participation from 214 teams across 82 universities.

02

Academic scholarship

₹3.4 lakh awarded for academic excellence over my first two years at Thapar.

03

Smart India Hackathon 2024

Selected among the top 40 of 190+ teams in the college's internal hackathon.

04 / KEEP IN TOUCH

Good questions.
New possibilities.

For an opportunity, a project, or a conversation about ML.