Mar 2026 —
Present
Deep Learning Compiler Engineer · NVIDIA
- Formal methods of verification for MLIR-based deep learning compilers.
- …and agents.
$ whoami samarth narang deep learning compiler engineer, nvidia new york, ny # previously qualcomm, degirum
$ cat work.txt nvidia 2026 — now formal verification qualcomm 2024 — 2026 hexagon npu, tiling degirum 2022 — 2024 accelerator backend # the long version is below
$ ls stack/ daily c++ python mlir llvm often pytorch assembly unix sometimes java sql flutter
$ tail -f now.log → making sure deep learning compilers do not break, mostly → writing at /tiled-thoughts → eating something, or in the gym so i can eat something
four commands, that's the site
about
I'm a deep learning compiler engineer at NVIDIA, working on formal methods of verification for MLIR-based deep learning compilers. Before that I spent two years at Qualcomm as one of the core contributors to a new MLIR compiler for AI inference on Hexagon NPUs.
My background and interests lie in compiler construction, machine learning, computer architecture and systems programming.
Outside of work I try (and mostly fail) to keep things balanced:

experience
Mar 2026 —
Present
Jan 2024 —
Mar 2026
Jun 2023 —
Jan 2024
ML compiler backend
ML deployment infrastructure
May 2022 —
May 2023
Machine learning team · Aug 2022 – May 2023
timm repository into DeGirum's model zoo.Embedded software team · Feb 2022 – Aug 2022
education
Aug 2024 —
Aug 2025
Aug 2020 —
May 2023
publications
projects
Ongoing contributor to the LLVM project. Patches across MLIR, Clang, LLVM optimization passes and Flang.
Artifact for the HASP 2026 paper. Matched Triton and CUDA implementations of the same SumCheck prover kernels over prime fields, plus the drivers that time and check them.
An RTL debugging loop on top of CHIA that verifies a patch independently, then checks the agent's diagnosis against a simulation-derived first-divergence label. Artifact for Right Fix, Wrong Reason: Repair and Diagnosis in a CHIA RTL Debugging Loop, accepted at the A³ Workshop CHIA hackathon, MICRO 2026.
An experimental MLIR dialect for high-frequency trading strategies, lowered end to end into a shared library a C++ engine runs against a toy market.
tick.on_booktick.order.sendtick.order.canceltick.risk.*earlier · mobile
A social app for recording memories for family, on a serverless AWS backend.
Price alerts for crypto traders, wired into exchange APIs for real-time data. App Store
writing
Compiler passes, ML systems, and the occasional detour. Long-form notes from the path between code and silicon.
A practical mental-model bridge from LLVM IR to MLIR.
Hand-written NEON intrinsics, loop unrolling, and a spoiler you can probably guess.
Everyone agrees it matters; the details stay fuzzy. Here are the details.
Collectives, interconnects, and the communication patterns behind efficient distributed training.
Zero-copy views and safe mutation loops for faster LLVM passes.
skills
contact
Reach out to talk MLIR, compilers, architecture, etc. etc.
New York, NY