Machine Learning Tech Lead
Oman,
Oman
Job Purpose
Rihal's Machine Learning Department delivers awarded client scopes across GenAI solutions, agentic systems, document intelligence, computer vision, and ML-powered products, including on-premise LLM deployments on Rihal's GPU infrastructure. The ML Tech Lead reports to the Machine Learning Manager and directly manages a defined group of ML engineers and their projects. The role combines people leadership with strong hands-on technical oversight to ensure high-quality delivery, sound engineering standards, and consistent ownership and commitment across the assigned team.
Duties and Responsibilities
- Technical Delivery & Client Project Execution
- Oversee day-to-day technical delivery for the ML engineers and client projects within the Tech Lead's assigned scope.
- Oversee end-to-end ML delivery: data readiness, model/LLM selection, evaluation, deployment, and monitoring.
- Ensure the team understands client requirements, acceptance criteria, technical constraints, priorities, and delivery commitments.
- Ensure evaluation criteria and acceptance metrics (e.g., accuracy, latency, cost-per-token) are agreed with clients before build.
- Guide and review the team's solution designs, technical plans, estimates, implementation approach, and delivery readiness.
- Monitor progress, blockers, risks, defects, and dependencies; resolve issues or escalate early to the Machine Learning Manager.
2. Technical Standards & Engineering Quality
- Enforce Rihal's engineering standards across prompt/agent design, evaluation harnesses, model versioning, data handling, MLOps practices, coding, architecture, code review, testing, security, documentation, and deployment.
- Review significant technical decisions (including model choice, fine-tuning vs. RAG vs. agentic approaches, and GPU resource sizing) and ensure solutions are secure, scalable, reliable, maintainable, and appropriate for the client environment.
- Ensure responsible AI practices across projects, including data privacy, hallucination mitigation, and output guardrails, particularly for government and regulated clients.
- Maintain effective design reviews, peer reviews, code reviews, technical validation, and quality-control practices across assigned projects.
- Identify recurring defects, rework, technical debt, and engineering weaknesses and ensure corrective action is taken.
3. Team Leadership, Ownership & Accountability
- Directly manage the assigned ML engineers, setting clear priorities, responsibilities, expectations, and technical direction.
- Build a culture of ownership, commitment, accountability, professionalism, and follow-through in daily delivery.
- Ensure engineers communicate blockers early, meet agreed commitments, and remain accountable for the quality of their outputs.
- Provide timely coaching and feedback, and escalate material performance, conduct, capacity, or capability concerns to the Machine Learning Manager
4. Technical Competency & Mentorship
- Maintain strong hands-on ML engineering competency and sufficient technical depth to guide, review, troubleshoot, and challenge the team's work.
- Act as the technical escalation point for inference serving (e.g., vLLM/SGLang), GPU cluster issues, pipeline failures, production model degradation, integration, debugging, and deployment issues within the assigned scope.
- Mentor engineers, identify skill gaps, and work with the Machine Learning Manager on targeted development, staffing, or hiring actions.
5. Resource, Client & Stakeholder Coordination
- Plan engineer allocation and GPU capacity within the assigned subset based on project priorities, skills, availability, and delivery needs.
- Maintain visibility of workload, capacity, project health, GPU utilization, and resource risks, and raise concerns early with the Machine Learning Manager.
- Represent the assigned ML team in client technical workshops, delivery meetings, project reviews, and technical escalations as required.
- Coordinate effectively with project managers, QA, DevOps, infrastructure, cybersecurity, and other delivery functions to resolve technical issues, including GPU cluster scheduling and resource allocation.
6. Delivery Performance & Continuous Improvement
- Track technical quality and delivery health across assigned projects, including defects, rework, technical debt, predictability, model performance drift, inference costs, and client-impacting issues.
- Provide the Machine Learning Manager with clear updates on project status, team performance, ownership, capacity, risks, and required actions.
- Drive lessons learned and continuous improvement in engineering discipline, estimation, evaluation practices, review practices, documentation, and delivery consistency
Requirements
ML Engineering & Technical Leadership
- Strong hands-on experience deploying and serving LLMs on dedicated GPU infrastructure (e.g., vLLM, SGLang, or TGI) in production environments.
- Demonstrated experience building agentic systems or GenAI applications end-to-end with genuine ownership of technical decisions.
- Strong proficiency in Python, Kubernetes/containerization, and inference optimization (e.g., quantization, batching, KV cache management).
- Strong troubleshooting and problem-solving ability across production ML infrastructure, including memory, throughput, and serving issues.
- Solid understanding of evaluation methodology for LLM and ML systems, including designing evaluation harnesses and acceptance metrics.
Delivery, Quality & People Leadership
- Demonstrated experience leading technical delivery of client-facing ML/AI projects while maintaining quality and delivery discipline.
- Strong understanding of estimation, agile delivery, testing, code review, secure development, release management, and technical documentation.
- Experience leading and mentoring ML engineers while remaining technically hands-on and accountable for team outcomes.
- Ability to set expectations, delegate, follow up on commitments, address underperformance, and build a culture of ownership and continuous improvement.
Client & Communication
- Comfortable working directly with client technical and business stakeholders throughout delivery, review, escalation, acceptance, and handover
- Ability to communicate technical decisions, risks, trade-offs, progress, and required actions clearly to both technical and non-technical stakeholders, including explaining model behavior and limitations to non-technical audiences.
Qualifications & Experience
- Bachelor's degree in Computer Science, Software Engineering, Information Technology, Data Science, or a related quantitative field, or equivalent practical experience.
- 4+ years of ML/AI engineering experience, including demonstrated technical leadership of ML engineers or ML project delivery.
- Demonstrated track record of delivering production ML/GenAI systems and leading engineers across client-facing projects or complex AI initiatives.
- Experience in software services, consulting, systems integration, or project-based delivery is an advantage.