college station, tx

Leonardo Schaub

I build agent infrastructure and study how large language models actually run underneath. Most recently I rebuilt Qualcomm's NOVA observability agent into a multi-agent swarm; right now I'm researching inference-time optimization at Texas A&M.

education
B.S. Computer Science, Texas A&M University — May 2028
minors
Math, Japanese
May 2026 Aug 2026

Software Engineering Intern — Qualcomm Inc.

San Diego, CA

NOVA (Next-Generation Observability Virtual Agent)

  • Migrated NOVA from a monolithic agent to a multi-agent Swarm architecture — 86% lower token usage, 72% faster responses; prototyped Supervisor and Supervisor/Swarm hybrid designs.
  • Designed concurrent tool-result compression and de-duplication in LangGraph, cutting response time ~46%.
  • Built an agent-regression benchmarking framework using golden test cases and LLM-as-a-judge rubrics to compare accuracy, latency, and token efficiency across releases.
May 2025 Aug 2025

Machine Learning Engineering Intern — Advantest America Inc.

Austin, TX
  • Reduced SmartTest 7/8 graph generation time by 80%+ with an NLP-driven interface translating natural language into automated SQL queries and visualizations.
  • Designed agent orchestration infrastructure using MCP servers to safely execute external tools and enforce structured workflows across LLM systems.
  • Automated updates to the Test Documentation Center by generating relevance reports for RAG retraining, cutting manual document review time 60%+.
Mar 2026 Present

Undergraduate ML Researcher — Texas A&M University

College Station, TX
  • Researching LLM inference and architecture optimization — speculative decoding, quantization, KV-cache optimization, and batching — to reduce latency, memory, and compute cost.
  • Investigating Mixture-of-Experts, dense-to-MoE transformations, and multi-agent architectures to study tradeoffs between capacity, routing, and inference efficiency.
  • Experimenting with abliteration and representation-level interventions to analyze how targeted weight/activation changes affect model behavior.
  • Building a reproducible PyTorch inference harness benchmarking TTFT, inter-token latency, tokens/sec, and peak memory.

LoL ML Pipeline

2026

End-to-end ML pipeline predicting professional League of Legends match outcomes from pre-game signals only — team and player identity, draft picks and bans, side selection, patch context — deliberately excluding post-game statistics to eliminate label leakage.

  • Deterministic feature engineering over versioned raw match data with immutable dataset snapshots, so every training run is reproducible from a resolvable dataset identifier.
  • Experiment tracking, model artifact versioning, and drift monitoring hooks to support scheduled retraining — prioritizing system design correctness over short-term accuracy.
Python TensorFlow Pandas Docker AWS Snowflake
languages

Python, C++, Java, Haskell, SQL

ml / ai

TensorFlow, PyTorch, LangGraph, MCP, RAG, Transformers, LLM Inference Optimization

data / observability

NumPy, Pandas, PostgreSQL, Splunk, Datadog, Langfuse, Langsmith

cloud & devops

AWS (S3, Batch, ECR), Docker, Git, Linux, CI/CD Pipelines

leadership

Texas A&M College of Engineering Student Ambassador (Fall 2025–present) — leads engineering tours for prospective students, families, and VIPs, engaging ~400 participants weekly.

honors

Dean's Honor Roll (Spring 2025, Fall 2025) · Dean's Excellence Award (Spring 2026)