About this site

samarth.sh
$ whoami

samarth narang
deep learning compiler engineer, nvidia
new york, ny

# previously qualcomm, degirum

four commands, that's the site

about

What I do.

I'm a deep learning compiler engineer at NVIDIA, working on formal methods of verification for MLIR-based deep learning compilers. Before that I spent two years at Qualcomm as one of the core contributors to a new MLIR compiler for AI inference on Hexagon NPUs.

My background and interests lie in compiler construction, machine learning, computer architecture and systems programming.

Outside of work I try (and mostly fail) to keep things balanced:

  • 🍜 major foodie
  • 🏋️ regular at the gym
  • ✈️ always planning the next trip
  • ⚽ follows almost every sport
  • ♟️ former competitive chess player
Samarth Narang

experience

Work.

Mar 2026 —
Present

Deep Learning Compiler Engineer · NVIDIA

  • Formal methods of verification for MLIR-based deep learning compilers.
  • …and agents.
  • MLIR
  • formal verification
  • LLVM

Jan 2024 —
Mar 2026

Deep Learning Compiler Engineer · Qualcomm

  • One of the core contributors to a new MLIR-based compiler stack for Qualcomm's AI inference on Hexagon NPUs — from prototyping through benchmarking of the Triton-based compilation path. Open sourced at qualcomm/hexagon-mlir.
  • Extended a sequence of MLIR passes — multi-level tiling, fusion, vectorization, multithreading — culminating in lowering to LLVM for NPU backend codegen.
  • Worked on the tiling algorithms used by Qualcomm's proprietary Hexagon NPU compiler for on-device AI inference.
  • Handled critical performance issues hit by customers running LLMs on Qualcomm NSPs.
  • MLIR
  • Triton
  • tiling
  • fusion
  • vectorization
  • NPU

Jun 2023 —
Jan 2024

Machine Learning Engineer · DeGirum

ML compiler backend

  • Built out an AI compiler for DeGirum's hardware accelerator, focused on performance and widening the set of models it could compile.
  • Implemented SIMD parallelism, vector processing, and pipelining strategies to cut data movement and memory footprint.
  • Extended the compiler to support LLMs, working through the details of transformer architectures.
  • Benchmarked CPU cores on FPGA to decide what to offload during real-time model execution.

ML deployment infrastructure

  • Compiled models with the DeGirum compiler and integrated them into the DeGirum ecosystem via a Flask API.
  • Wrote the PyTest suites and CI/CD pipelines that automated deployment and testing.
  • SIMD
  • LLM support
  • FPGA
  • Flask
  • CI/CD

May 2022 —
May 2023

Software Engineering Co-op · DeGirum

Machine learning team · Aug 2022 – May 2023

  • Compiled PyTorch models onto company-specific compilers; handled FP32 → UInt8 quantization to widen model coverage.
  • Ported 132 models (quantized and float) from the timm repository into DeGirum's model zoo.

Embedded software team · Feb 2022 – Aug 2022

  • Built a UART interface monitor in RISC-V assembly as a field-engineer debug tool.
  • Wrote ROM code routines for MBIST operations and redesigned the existing MBIST tests to finish 40% faster.
  • Wrote embedded C tests for pre-silicon RTL validation and extended the Verilog test bench.
  • PyTorch
  • quantization
  • RISC-V
  • MBIST
  • Verilog

education

Degrees.

Aug 2024 —
Aug 2025

M.S. Computer Science · UT Austin

  • GPA 3.94. Coursework in compiler construction and implementation of programming languages, virtualization, NLP, reinforcement learning, generative AI, and optimization.

Aug 2020 —
May 2023

B.S. Computer Science & Mathematics · UMass Amherst

  • GPA 3.98 — Summa Cum Laude, Dean's List every semester, Chancellor's Scholarship ($16,000/yr).
  • Coursework in operating systems, networking, computer systems, algorithms and data structures, machine learning, and regression analysis.

publications

Papers.

projects

Personal time.

Microtick

MLIR · LLVM · C++

An experimental domain-specific language for high-frequency trading strategies, built on MLIR. A custom tick dialect plus an end-to-end lowering pipeline that takes strategy IR down to a shared library a C++ engine can dlopen and run against a toy market with fills, positions, and P&L.

tick.on_booktick.order.sendtick.order.canceltick.risk.*

LLVM & MLIR upstream

open source

Ongoing contributor to the LLVM project — patches across MLIR, Clang, LLVM optimization passes, and Flang.

earlier · mobile

  • CaptureMyHippo

    Flutter · Flask · AWS Lambda · DynamoDB

    A non-conventional social app for recording memories for family and loved ones, on a serverless AWS backend. App Store

  • NoFinishLine

    Flutter · Flask · AWS · Boto3

    iOS workout tracker over a searchable exercise database with levels, categories, and instructions, backed by a RESTful API. GitHub

  • CryptoPriceAlert

    Flutter · Flask · Firebase

    Price alerts for crypto traders, wired into major exchange APIs for real-time data. App Store

writing

Tiled Thoughts — a verbose debug build

Compiler passes, ML systems, and the occasional detour. Long-form notes from the path between code and silicon.

skills

Tools.

Proficient

  • C
  • C++
  • Python
  • MLIR
  • LLVM
  • PyTorch
  • Assembly
  • Git
  • Unix
  • Flutter

Familiar

  • Java
  • RISC-V ISA
  • SQL
  • TensorFlow

contact

Get in touch.

Reach out to talk MLIR, compilers, architecture, etc. etc.

New York, NY