Yuhao Cheng
A portfolio in ten scenesChampaign, IL

YuhaoCheng.

ML systems & research engineer — M.S. Computer Science, UIUC

Scroll to begin
SC 01Prologue — the whole set, from above

I build the systems underneath models — the pipelines that train them, serve them, and measure what they can actually do.

Seven sets below, one for each chapter: a campus, an evaluation pipeline, a GPU cluster, a dashboard studio, an inference farm, a benchmark gallery and a retrieval funnel.

0M+
LLM calls orchestrated in one evaluation run
0d → 21h
Full benchmark run, after load-testing the pipeline
0 img/min
SDXL throughput across 8 GPUs, up from 6
0+
WebArena tasks for SFT → RL web agents
SC 02Ext. Main Quad, Urbana-Champaign

Education

University of Illinois Urbana-Champaign

Aug 2025 — Dec 2026

Master of Computer Science

Aug 2021 — May 2025

B.S. in Computer Engineering

SC 03Int. Evaluation pipeline — 2026

Research

TIMAN Group, UIUC

Research Assistant · May 2026 — Sep 2026

Auditing large language models at scale.

  • Built a Python evaluation pipeline that executed 3M+ LLM calls across 8 models on a 64-worker pool, with a token-bucket rate limiter and exponential-backoff retries.
  • Load-tested the model endpoints to their limit, raising throughput from 40 to 2,400 calls/min — a full run went from an estimated 52 days to 21 hours.
  • Designed BDSU audits of GPT-4o, GPT-4o-mini and o1, decomposing behavior into demographic disparity, prompt sensitivity and generation uncertainty.
  • Made runs resumable with request-level caching and checkpointing; every generation, score and metric is tracked in Weights & Biases.
LLM EvaluationConcurrencyW&B
SC 04Int. GPU cluster — 4 nodes, 16 GPUs

Research

IBM–Illinois Discovery Accelerator Institute

Research Intern · Jul 2024 — Dec 2024

Teaching an 8B model to browse the web.

  • Built an SFT-to-RL pipeline for Llama-3.1-8B web agents over 800+ WebArena tasks: browser interaction, trajectory collection and automated evaluation.
  • Implemented actor-critic RL with dense trajectory rewards — task success, progress, URL similarity, exploration bonuses, state-change penalties.
  • Scaled full-parameter training across 4 nodes / 16 GPUs with PyTorch, Accelerate and DeepSpeed ZeRO-3; 16 parallel browser workers cut an eval pass from 20 h to 3 h.
RLLLM AgentsDeepSpeedSLURM
SC 05Int. Dashboard studio — summer 2025

Industry

visibilityx.ai

Frontend Developer Intern · Jun 2025 — Aug 2025

Data-heavy dashboards that load fast.

  • Built a Vue 3 + TypeScript SPA with 10+ data-intensive dashboard views; extracted 12 shared components and composables, removing ~2,000 lines of duplicated logic.
  • Cut median dashboard load from 2.4 s to 1.6 s with parallel REST calls, Pinia response caching and per-route lazy-loading of ECharts.
  • Wrote 180+ Vitest unit tests (80% line coverage), run in CI on every pull request.
Vue 3TypeScriptEChartsVitest
SC 06Int. Inference farm — SDXL on 8 GPUs

Industry

HiABR Lab

Backend Developer Intern · May 2024 — Aug 2024

Serving Stable Diffusion XL on a GPU cluster.

  • Deployed SDXL as a multi-node inference service — 2 nodes, 8 GPUs, one FastAPI worker per GPU — behind Nginx least-connections routing with health checks.
  • Traced OOM crashes to overlapping requests on 24 GB GPUs; fixed them with bounded per-GPU job queues, then added dynamic micro-batching (4 prompts / 50 ms window).
  • Under Locust load, went from failing above 4 concurrent users to 64 with zero errors; throughput scaled from 6 to 140 images/min at 21 s p95.
  • Built a URL-shortening service on FastAPI, PostgreSQL and Redis with sliding-window rate limiting, idempotency keys and Prometheus metrics.
FastAPIMulti-GPUNginxRedis
SC 07Int. Gallery — 25 of 27 tasks on the wall

EMNLP 2026

VGI-Bench

Probing Visual Intelligence in Video Generation Models

27tasks
810instances
20models
  • Co-developed a visual-reasoning benchmark across four task domains and seven capability dimensions.
  • Built pipelines automating inference, collection and analysis for 9 video and 11 image generation models; studied failure modes, input sensitivity and reasoning dynamics.
SC 08Int. Retrieval funnel — ongoing

2026 — ongoing

PIR-Arena

Proactive Information Recommendation Benchmark

34real scenes
831minutes
11need types
  • A multimodal benchmark for deciding when and what information to surface from continuous user context.
  • Retrieval pipeline linking LLM query generation with BM25, dense retrieval and cross-encoder reranking over 50k documents at 1.8 s p95; evaluated on precision, recall, relevance and timeliness.
SC 09Montage — the toolkit

01 Languages

  • Python
  • TypeScript
  • JavaScript
  • C / C++
  • SQL
  • Java
  • Bash

02 ML Systems

  • PyTorch
  • DeepSpeed
  • Accelerate
  • SLURM
  • Weights & Biases

03 Backend & Infra

  • FastAPI
  • PostgreSQL
  • Redis
  • Docker
  • Nginx
  • Prometheus
  • Locust
  • AWS
  • GCP

04 Frontend

  • React
  • Vue 3
  • Node.js
  • Express
  • ECharts
  • Vitest
SC 10Fade out

Let’s make the next scene together.

Open to ML systems, infrastructure and research engineering roles — from Dec 2026.

yuhaoc7@outlook.com