Skills
About the Role
Intel is building agentic AI that combines local intelligence with cloud capabilities—so models can run efficiently at the edge while preserving user privacy and reducing token costs. As a Senior Inference Optimization Engineer, you’ll help optimize how AI inference performs across local and edge runtimes.
Responsibilities
- Improve inference performance for small, efficient models running on user devices and edge/on-prem environments
- Design, implement, and validate optimization strategies for local/edge runtimes
- Collaborate with ML, systems, and engineering teams to profile bottlenecks and reduce latency and compute overhead
- Develop and maintain benchmarking and evaluation workflows to measure end-to-end inference quality and speed
- Partner with stakeholders to translate performance requirements into practical engineering solutions
Requirements
- Experience optimizing model inference performance in production or near-production environments
- Strong understanding of runtime behavior (e.g., memory, batching, quantization, and execution efficiency)
- Proficiency with performance profiling and tuning tools
- Experience working with ML inference stacks and deployment/runtime environments
- Excellent problem-solving skills and ability to collaborate across disciplines
Benefits
- Opportunity to work on privacy-focused, resource-efficient AI systems
- Collaborate with teams building local + cloud agentic AI
- Contribute to technologies that improve real-world model performance