Staff Software Engineer, AI Inference
Syllo
United States · Posted Jul 13
Job description
Responsibilities: Lead the design and development of a production AI inference platform to serve large language models. Define the technical roadmap for model serving, runtime optimization, and infrastructure scalability.
Requirements: Requires significant experience building production LLM serving infrastructure and deep knowledge of modern inference runtimes. Proficiency in Python and a systems programming language like Go, Rust, or C++ is essential.
Key skills: AI Inference, LLM Serving, Distributed Systems, GPU Optimization, Python, Go, Rust, C++, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, Backend Infrastructure, Technical Architecture
Keywords: AI Inference, LLM, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, GPU Utilization, Latency Optimization, Throughput, Distributed Systems, Python, Go, Rust, C++, Model Deployment, Autoscaling, Observability, Agentic AI, Legal Tech, Infrastructure Engineering, Runtime Optimization, Benchmarking, Production Systems