Logo NVIDIA

LLM Inference Engineer

NVIDIA· Austin, TX· Ref. EXP-2026-0019 4.5 (5,400 reviews)
Easy applyFull-timeHybrid
View without an account · free sign-in only to apply
Salary
$155k – $220k / year
Type
Full-time
Location
Austin, TX · Austin (Texas) · Hybrid
Valid through
10/31/2026

Additional details

Posted
Posted 4 days ago
Applications
31
on Ganloss
Experience
Mid · 3+ years
Education
BS/MS CS, ML, or equivalent experience
Region
Texas
Team
NVIDIA — AI team
Company size
25000+
Founded
1993

Job description

NVIDIA is hiring a LLM Inference Engineer (Full-time) hybrid. This llm engineer role is part of the AI hiring market in Austin (Texas). Ship high-throughput LLM serving stacks for enterprise NIM deployments. Listed compensation: $155k – $220k / year. Stack: TensorRT-LLM, vLLM, Python, GPU.

LLM Engineer profiles are in high demand in Austin: Austin’s tech scene keeps growing with AI product, robotics, and data roles at scale-ups and enterprise R&D labs.…

The team

NVIDIA — AI team

Responsibilities

  • Ship production AI features with measurable quality and latency targets
  • Partner with product and design on roadmap and customer feedback
  • Improve observability, evals, and safe rollout practices

Requirements

  • Strong Python or TypeScript fundamentals and 3+ years relevant experience
  • Hands-on LLM, ML, or data platform work in production
  • Clear written communication and async collaboration

Tech stack

TensorRT-LLMvLLMPythonGPU

Benefits

Health, dental, vision
401(k) match
Equity
Learning budget
Flexible PTO

Hiring process

  1. 1Recruiter screen (30 min)
  2. 2Technical interview (60 min)
  3. 3Team fit
  4. 4Offer

Experience: Mid · 3+ years · Education: BS/MS CS, ML, or equivalent experience

Salary guide — LLM Engineer in Austin

Go further

Prepare your application with our AI career guides — role sheets, resume, interview, and salary benchmarks.

FAQ — LLM Inference Engineer

What is the salary for LLM Inference Engineer at NVIDIA?
Listed range: $155k – $220k / year (Full-time). Packages often include PTO, remote options, and equity depending on the company.
Where is the LLM Inference Engineer role based?
Austin, TX, Austin (Texas). Work mode: Hybrid.
How do I apply to NVIDIA?
Create a free Ganloss candidate account, then apply in one click. Job reference: EXP-2026-0019. Process: Recruiter screen (30 min) → Technical interview (60 min) → Team fit → Offer.
Is NVIDIA hiring other AI roles?
NVIDIA (Hardware & AI) hires across LLM, ML, data, and AI product. See all openings on the Ganloss company page.

Apply in one click

Typical reply within 48h

Your profile is sent directly to NVIDIA's hiring team.

Apply to this job

Create a free candidate account (30 sec) to send your profile to NVIDIA. The job page stays public without signing up.