Staff Software Engineer, AI Inference

Syllo

United States · Posted Jul 13


Job description

Responsibilities: Lead the design and development of a production AI inference platform to serve large language models. Define the technical roadmap for model serving, runtime optimization, and infrastructure scalability.

Requirements: Requires significant experience building production LLM serving infrastructure and deep knowledge of modern inference runtimes. Proficiency in Python and a systems programming language like Go, Rust, or C++ is essential.

Key skills: AI Inference, LLM Serving, Distributed Systems, GPU Optimization, Python, Go, Rust, C++, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, Backend Infrastructure, Technical Architecture

Keywords: AI Inference, LLM, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, GPU Utilization, Latency Optimization, Throughput, Distributed Systems, Python, Go, Rust, C++, Model Deployment, Autoscaling, Observability, Agentic AI, Legal Tech, Infrastructure Engineering, Runtime Optimization, Benchmarking, Production Systems

Land this job faster with Remote Job Match

Free account: browse thousands of remote roles, no degree needed. Upgrade to tailor your resume to each job with AI and prep for the interview.

Create free account

← Browse more remote jobs