Text version

Loading the district…

The work

13 projects, one district.

Each building up in the town is one of these. Here they are up close — the research, the results, and the exact sources behind every claim. Scroll the town, or read them here.

The headline result. Retained signal R at fixed relative-depth layers across the decorrelation ladder, with 95% paired-bootstrap intervals and the frozen illusion / partial / abstract bands. Removing markers roughly halves retention at both scales.

Paper · resources/116_When_ReAct_Phase_Probes_Lo.pdf p.5, Figure 2

01

When ReAct Phase Probes Look Abstract

Mechanistic Interpretability ResearcherWork in progressApr 2026 – Present

A preregistered test of whether a linear probe that decodes ReAct reasoning phase at 98% accuracy is reading an abstract reasoning state or a correlated surface cue.

mechanistic-interpretabilityprobingsparse-autoencodersgemma-2preregistrationtransformerlens

The calibration stress test. A plain cross-fitted calibrated prior (Brier 0.216) beats the grounded LLM hybrid (0.235) — the result that fails the predictive-validity gate.

Paper · resources/36_More_Context_Is_Not_Validat (1).pdf p.3, Table 2

02

More Context Is Not Validation

AI Researcher Intern, Production LLM Systems & Evaluation InfrastructureWork in progressMay 2026 – Present

A three-gate audit of whether LLM agents that simulate individual listeners actually predict them, or only sound like they do.

llm-evaluationcalibrationagent-simulationbaselinesconstruct-validitybootstrap
ME
03

Medical Policy Digital Twin

Tech Projects Lead, Technology Partner (4 concurrent tech companies) — engineering lead on this projectOTCR ConsultingAug 2025 – Present

A policy simulator that replays a payer’s historical claims against proposed policy edits and forecasts the cost impact before the policy ships.

healthcaresimulationetlpostgresqlawsstreamlit

Project poster — NCSA Research Conference 2026. Click to read it full size.

Source · POSTER (NCSA, Spring 2026) — Predicting Internal Migration in the Brazilian Amazon

04

Predicting Internal Migration in the Brazilian Amazon

ML Lead & SPIN Research Fellow, Distributed Training & Data InfrastructureNCSA Research Conference 2026 (poster)Aug 2025 – Present

A graph neural network that predicts internal migration in the Brazilian Amazon as directed flows between municipality pairs — paired with gradient-boosted trees and SHAP for feature-level explanation across 5,570 municipalities and 300K+ town-to-town flows.

graph-neural-networksgraphsage-attentionxgboostshapgeospatialmigration

Tier 2. The probe hits ~93% at every layer against a 25% baseline. The README reads that flatness as a warning sign, not a finding.

Repository · github.com/Chinmayrawat15/ReAct_interp — figures/tier2_probe_accuracy.png

05

ReAct Interp — the GPT-2 pilot

Sole committer on the public repositoryApr 2026 (repository created and last pushed 2026-04-28)

A one-week interpretability sprint on GPT-2 Small that found ReAct phase trivially decodable at every layer — and said so, listing the controls it had not run.

mechanistic-interpretabilitygpt-2sparse-autoencodersprobingexploratorypython
RepositoryREADME

Figure 1 from the paper — mean temperature, precipitation, and GDP per capita across Mexican municipalities. Click to enlarge.

Source · PAPER-C p.2, Figure 1 — Beyond Gravity Models

06

Beyond Gravity Models

Co-authorUniversity of Illinois Urbana-Champaign (NSF REU)2026

A co-authored paper on a graph neural network for internal migration in Mexico that beats gravity-model baselines by adding a spatial gate and an income-conditioned climate gate — with a built-in interpretability boost.

graph-neural-networksmigrationgravity-modelsspatial-lagclimateinterpretability

Two of the upstream projects contributed to. Logos are the trademarks of their respective projects and appear here to identify them, not to imply endorsement.

Source · Project logos (provided) — OpenClaw and PyTorch

07

Open Source Contributions

Open Source Contributor, ML frameworks & agent toolingJan 2026 – Sep 2026 (dates of first and most recent pull requests found)

Merged bug fixes to TransformerLens, marimo, OpenClaw and pylint, plus open pull requests to PyTorch, LiteLLM, Prefect, vLLM and Hermes.

open-sourcetransformerlenspytorchmarimoopenclawlitellm

The live Phantom demo. Intended for authorised testing of targets you own or have permission to test.

Web · https://phantom-hackillinois.pages.dev (live site, retrieved 2026-08-26)

08

Phantom

Team Lead (team HackVoyagers)HackIllinois 2026 · YC Summer School 20262026 (repository commits Feb–Mar 2026)

An autonomous penetration-testing agent: point it at a target and it crawls the attack surface, plans and executes multi-step exploit chains in a ReAct loop, and returns CVSS-scored findings with reproduction steps in under 60 seconds.

securityagentsreact-loopplaywrightcloudflare-workersmodal
Live demo

PocketFM’s public "Copilot" writing product, which advertises adapting stories to 10+ languages.

Source · PocketFM — Copilot product marketing image ("COPILOT by PocketFM")

09

Adaptation

AI Researcher Intern, Production LLM Systems & Evaluation InfrastructurePocketFMMay 2026 – Present

A production LLM service that rewrites full scripts into a target language, adapting cultural references rather than translating them.

llm-applicationlocalizationproductionnlp

PocketFM’s audio-series app — the catalogue whose launch outcomes the Blockbuster Engine is built to forecast. Illustrative product image, not the engine itself.

Source · PocketFM — app / product marketing image

10

Blockbuster Engine

Software Engineer Intern, GenAIPocketFM (Bangalore, India)May 2025 – Jul 2025

An ensemble of fine-tuned LLMs and tree/neural models predicting content launch outcomes, deployed as containerised microservices.

loraensemblesgradient-boostingmicroservicesdockerproduction-ml

BeClear — brand wordmark. Logo supplied by the founder; no public product page was found.

Source · BeClear — brand logo (provided)

11

BeClear

Co-Founder2024 – 2025

An AI college-admissions agent, built on a RAG pipeline with LoRA fine-tuning; sold in 2025 to a group of schools in India.

raglorastartupedtechllm-application
MU
12

Multi-Agent Systems & Security

Researcher, Multi-Agent Systems & SecurityAlta Research Group, UIUC (advisor: Prof. Hao Peng)Aug 2026 – Present

Two connected questions in multi-agent coding safety, with Prof. Hao Peng: whether sub-agents can covertly pass confidential data past an LLM monitor, and whether the reasoning agents write down is what actually drives what they do.

multi-agentai-safetysecurityevaluationllm-monitoringreasoning-faithfulness
TE
13

Test-Time Scaling for Coding Agents

Researcher — Agentic Systems (with Prof. Hao Peng)Alta Research Group, UIUC (advisor: Prof. Hao Peng)Sep 2026 – Present

An early system that lets coding agents explore a codebase over many interactions to surface new, reusable knowledge about a task — with exploration pushed into the background so it stays cheap for users.

agentscoding-agentstest-time-scalingmulti-agentorchestrationearly-stage