This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Inference Engineer based in Brazil. This is an opportunity to become the first dedicated engineer responsible for building and owning an inference platform from the ground up.Youβll work closely with the CTO to turn large language models into reliable, production-grade query-to-response systems running at scale.The role combines hands-on infrastructure engineering with performance optimization across GPU-based environments.Youβll work with modern serving technologies such as vLLM, SGLang, and TensorRT-LLM while solving real-world challenges around latency, cost, and throughput.As the platform matures, youβll take increasing ownership of its technical direction and collaborate with Product on the future inference roadmap.Youβll join a remote-first, open-source-oriented startup where engineering ownership, speed, and measurable customer impact are highly valued.This role is ideal for a senior engineer who wants significant autonomy and the opportunity to define how production inference infrastructure is built and scaled.