About this site

samarth.sh
$ whoami

samarth narang
deep learning compiler engineer, nvidia
new york, ny

# previously qualcomm, degirum

four commands, that's the site

about

What I do.

I'm a deep learning compiler engineer at NVIDIA, working on formal methods of verification for MLIR-based deep learning compilers. Before that I spent two years at Qualcomm as one of the core contributors to a new MLIR compiler for AI inference on Hexagon NPUs.

My background and interests lie in compiler construction, machine learning, computer architecture and systems programming.

Outside of work I try (and mostly fail) to keep things balanced:

  • 🍜 major foodie
  • 🏋️ regular at the gym
  • ✈️ always planning the next trip
  • ⚽ follows almost every sport
  • ♟️ former competitive chess player
Samarth Narang

experience

Work.

Mar 2026 —
Present

Deep Learning Compiler Engineer · NVIDIA

  • Formal methods of verification for MLIR-based deep learning compilers.
  • …and agents.
  • MLIR
  • formal verification
  • LLVM

Jan 2024 —
Mar 2026

Deep Learning Compiler Engineer · Qualcomm

  • One of the core contributors to a new MLIR-based compiler stack for Qualcomm's AI inference on Hexagon NPUs — from prototyping through benchmarking of the Triton-based compilation path. Open sourced at qualcomm/hexagon-mlir.
  • Extended a sequence of MLIR passes — multi-level tiling, fusion, vectorization, multithreading — culminating in lowering to LLVM for NPU backend codegen.
  • Worked on the tiling algorithms used by Qualcomm's proprietary Hexagon NPU compiler for on-device AI inference.
  • Handled critical performance issues hit by customers running LLMs on Qualcomm NSPs.
  • MLIR
  • Triton
  • tiling
  • fusion
  • vectorization
  • NPU

Jun 2023 —
Jan 2024

Machine Learning Engineer · DeGirum

ML compiler backend

  • Built out an AI compiler for DeGirum's hardware accelerator, focused on performance and widening the set of models it could compile.
  • Implemented SIMD parallelism, vector processing, and pipelining strategies to cut data movement and memory footprint.
  • Extended the compiler to support LLMs, working through the details of transformer architectures.
  • Benchmarked CPU cores on FPGA to decide what to offload during real-time model execution.

ML deployment infrastructure

  • Compiled models with the DeGirum compiler and integrated them into the DeGirum ecosystem via a Flask API.
  • Wrote the PyTest suites and CI/CD pipelines that automated deployment and testing.
  • SIMD
  • LLM support
  • FPGA
  • Flask
  • CI/CD

May 2022 —
May 2023

Software Engineering Co-op · DeGirum

Machine learning team · Aug 2022 – May 2023

  • Compiled PyTorch models onto company-specific compilers; handled FP32 → UInt8 quantization to widen model coverage.
  • Ported 132 models (quantized and float) from the timm repository into DeGirum's model zoo.

Embedded software team · Feb 2022 – Aug 2022

  • Built a UART interface monitor in RISC-V assembly as a field-engineer debug tool.
  • Wrote ROM code routines for MBIST operations and redesigned the existing MBIST tests to finish 40% faster.
  • Wrote embedded C tests for pre-silicon RTL validation and extended the Verilog test bench.
  • PyTorch
  • quantization
  • RISC-V
  • MBIST
  • Verilog

education

Degrees.

Aug 2024 —
Aug 2025

M.S. Computer Science · UT Austin

  • GPA 3.94. Coursework in compiler construction and implementation of programming languages, virtualization, NLP, reinforcement learning, generative AI, and optimization.

Aug 2020 —
May 2023

B.S. Computer Science & Mathematics · UMass Amherst

  • GPA 3.98 — Summa Cum Laude, Dean's List every semester, Chancellor's Scholarship ($16,000/yr).
  • Coursework in operating systems, networking, computer systems, algorithms and data structures, machine learning, and regression analysis.

publications

Papers.

projects

Personal time.

LLVM & MLIR upstream

open source

Ongoing contributor to the LLVM project. Patches across MLIR, Clang, LLVM optimization passes and Flang.

zkDuel

Triton · CUDA · Python

Artifact for the HASP 2026 paper. Matched Triton and CUDA implementations of the same SumCheck prover kernels over prime fields, plus the drivers that time and check them.

YARROW

Verilog · Python

An RTL debugging loop on top of CHIA that verifies a patch independently, then checks the agent's diagnosis against a simulation-derived first-divergence label. Artifact for Right Fix, Wrong Reason: Repair and Diagnosis in a CHIA RTL Debugging Loop, accepted at the A³ Workshop CHIA hackathon, MICRO 2026.

Microtick

MLIR · LLVM · C++

An experimental MLIR dialect for high-frequency trading strategies, lowered end to end into a shared library a C++ engine runs against a toy market.

tick.on_booktick.order.sendtick.order.canceltick.risk.*

earlier · mobile

  • CaptureMyHippo

    Flutter · Flask · AWS Lambda · DynamoDB

    A social app for recording memories for family, on a serverless AWS backend.

  • CryptoPriceAlert

    Flutter · Flask · Firebase

    Price alerts for crypto traders, wired into exchange APIs for real-time data. App Store

writing

Tiled Thoughts — a verbose debug build

Compiler passes, ML systems, and the occasional detour. Long-form notes from the path between code and silicon.

skills

Tools.

Proficient

  • C
  • C++
  • Python
  • MLIR
  • LLVM
  • PyTorch
  • Assembly
  • Git
  • Unix
  • Flutter

Familiar

  • Java
  • RISC-V ISA
  • SQL
  • TensorFlow

contact

Get in touch.

Reach out to talk MLIR, compilers, architecture, etc. etc.

New York, NY