Staff Ai Infrastructure Engineer

 

Description:

Staff AI Infrastructure Engineer, You’ll Be At The Core Of Building And Scaling The Technical Infrastructure For AI/ML Systems. You Will
 

  • Build reusable CI/CD workflows for model training, evaluation, and deployment — integrating Langfuse, GitHub Actions, and experiment tracking, etc.
  • Automate model versioning, approval workflows, and compliance checks across environments.
  • Build out a modular and scalable AI infrastructure stack — including vector databases, feature stores, model registries, and observability tooling.
  • Partner with engineering and data science to embed AI models and agents into real-time applications and workflows.
  • Continuously evaluate and integrate state-of-the-art AI tools (e.g. LangChain, LlamaIndex, vLLM, MLflow, BentoML, etc.).
  • Drive AI reliability and governance, enabling experimentation while ensuring compliance, security, and uptime.
  • Build and enhance AI/ML Model Performance
  • Ensure data accuracy, consistency and reliability, leading to better model training and inferencing
  • Deploy infrastructure to support offline and online evaluation of LLMs and agents — including regression testing, cost monitoring, and human-in-the-loop workflows.
  • Enable researchers to iterate quickly by providing sandboxes, dashboards, and reproducible environments.
     

What We’re Looking For
 

  • Write high-quality, maintainable software — primarily in Python, but we value engineering ability over language familiarity.
  • Have a strong background in scalable infrastructure, including:
    • Containerization and orchestration (e.g. Docker, Kubernetes)
    • Infrastructure-as-code and deployment (e.g. Terraform, CI/CD pipelines)
    • Monitoring and logging frameworks (e.g. Datadog, Prometheus, OpenTelemetry)
  • Understand and implement ML Ops best practices, including:
    • Model versioning and rollback strategies
    • Automated evaluation and drift detection
    • Scalable model and agent serving infrastructure (e.g. vLLM, Triton, BentoML)
  • Deploy and maintain LLM and agentic workflows in production, including:
    • Monitoring cost, latency, and performance
    • Capturing traces for analysis and debugging
    • Optimizing prompt/response flows with real-time data access
  • Demonstrate strong ownership and pragmatism, balancing infrastructure elegance with iterative delivery and measurable impact.

Learn About TRM Speed In This Position
 

  • Rapid Issue Resolution. TRM Engineers identify and resolve critical onsite issues in minutes to hours, not weeks. We create virtual war rooms, implement fixes, and share lessons with both customer stakeholders and internal teams within 48 hours.
  • Navigating Bureaucracy. We anticipate and address procedural hurdles, build trust with key stakeholders, and find alternative pathways to approvals. This keeps projects moving even in complex environments.
  • Efficient Knowledge Transfer. Engineers document and share updates in real time, ensuring the entire team—onsite and remote—has full visibility into plans, blockers, and resolutions. Knowledge sharing sessions and clear documentation reduce friction and accelerate delivery.

Organization TRM Labs
Industry IT / Telecom / Software Jobs
Occupational Category Staff AI Infrastructure Engineer
Job Location New York,USA
Shift Type Morning
Job Type Full Time
Gender No Preference
Career Level Intermediate
Experience 2 Years
Posted at 2026-07-19 6:52 pm
Expires on 2026-09-02