Engineering & Technology
Lead
Bulgaria; Poland

Lead Nvidia AI Consultant

About the role

In this role, you will combine deep hands-on engineering expertise in GPU inference, LLM deployment, and modern AI infrastructure with the consulting gravitas to lead client engagements, shape enterprise architecture decisions, and drive pre-sales. This is not a pure advisory role — you will own technical outcomes, drive AI adoption, and shape go-to-market solutions working closely with clients, Nvidia stakeholders, and internal engineering teams.

Responsibilities

  • Own and lead AI engagements end-to-end — from discovery, assessment, and strategy definition through architecture design and production implementation, including clearly scoped implementation plans and delivery outcomes
  • Translate complex business and operational challenges into AI use-case definitions, solution roadmaps, and corresponding reference architectures aligned with client scope and strategic objectives
  • Design and validate production-grade AI architectures leveraging NVIDIA technologies including NeMo, NIM, Triton, Riva, DeepStream, Metropolis, and Omniverse across cloud and on-premises environments
  • Define reference architectures for GenAI and Agentic AI solutions, advising on GPU-accelerated infrastructure design, Kubernetes-based workload orchestration, and MIG/vGPU partitioning strategies
  • Profile and benchmark AI inference deployments to validate GPU utilization, memory footprint, and cost-efficiency against target SLAs — not just functional correctness
  • Continuously tune deployment configurations — batch size, concurrency, tensor and pipeline parallelism, quantization level, and KV-cache settings — based on profiling data to hit optimal latency, throughput, and cost tradeoffs
  • Benchmark LLM serving stacks across frameworks such as vLLM, TGI, and Triton+TensorRT-LLM, measuring throughput, TTFT, TPOT, and tokens/sec/GPU to drive informed infrastructure decisions
  • Lead pre-sales activities including proposals, solution positioning, technical discovery sessions, and customer workshops, and drive GenAI proof-of-concept initiatives that demonstrate clear and measurable business value
  • Contribute to NVIDIA alliance go-to-market strategy by shaping industry offerings, reusable accelerators, demos, and solution blueprints
  • Drive thought leadership through whitepapers, technical blogs, conference presentations, and industry events while actively mentoring engineering and consulting teams on NVIDIA ecosystem technologies and modern AI/ML practices

Requirements

  • 6+ years of experience in AI consulting, GenAI/Agentic AI development, Machine Learning, or Deep Learning with a demonstrable track record of leading client-facing engagements end-to-end
  • Strong hands-on expertise in Generative AI, Agentic AI, multimodal AI, transformers, LLMs, and VLMs with practical experience across the full AI lifecycle including experimentation, fine-tuning, optimization, deployment, inference, and monitoring
  • Hands-on experience with Python and modern AI/ML frameworks including PyTorch, TensorFlow, Pandas, NumPy, and Hugging Face
  • Substantial hands-on experience with at least three NVIDIA ecosystem platforms including NeMo, NIM, Riva, Metropolis, Omniverse, Triton, DeepStream, or TensorRT-LLM
  • Hands-on experience deploying AI/ML workloads on Kubernetes including Helm, Operators, GPU device plugins, NVIDIA GPU Operator, and MIG/vGPU partitioning, alongside practical experience with self-hosted inference stacks such as vLLM or Ollama
  • Working knowledge of model quantization techniques and inference optimization with hands-on experience using GPU profiling tools including NVIDIA Nsight Systems, Nsight Compute, nvidia-smi, DCGM metrics, PyTorch Profiler, and vLLM/Triton metrics endpoints
  • Demonstrated ability to read GPU utilization signals, diagnose compute-bound, memory-bound, or I/O-bound bottlenecks, and translate findings into actionable architecture or configuration changes
  • Experience designing and deploying AI solutions in at least one major cloud environment — AWS, Azure, or GCP — with strong understanding of TensorRT, Triton Inference Server, CUDA, DeepStream, and ONNX
  • Strong advisory and stakeholder management capabilities with experience presenting to executive and leadership audiences and leading pre-sales engagements including proposals, technical discovery sessions, workshops, and solution positioning
  • Deep understanding of enterprise architecture, distributed systems, Big Data, SDLC, MLOps, and AI governance practices combined with strong communication, analytical, and problem-solving skills

SoftServe is an equal opportunity employer. Qualified applicants will receive consideration regardless of race, color, ancestry, ethnicity, national origin, religion, sex, sexual orientation, gender identity or expression, age, citizenship, disability, health condition, marital or family status, veteran status, or any other characteristic protected by applicable law.

#LI-Remote

Role Summary

Location

Bulgaria, Poland

Work type

Remote/Office

Direction

Engineering & Technology

Subdirection

Data Science

Tech level

Lead

Personal recruiter:

Olha Valchuk

Apply Now

Fill out the form, and we'll be in touch shortly.

CV/Resume in English will speed up its processing time
Upload file

About us

We are a digital engineering and technology consulting company where expertise grows alongside people. For more than 30 years, we have been elevating technology: helping organizations navigate complex business challenges by combining deep engineering knowledge with thoughtful, research-backed innovation. Our teams work across key areas: digital engineering, data and analytics, Сloud, and AI/ML. In each, we deliver practical, scalable solutions rooted in real business needs and measurable human impact.

You bring your perspective and ambition. We create an environment where your work meets clarity, confidence, and purpose.

About us

We offer

Flexible Work Model

Work from home, from the office, or in a hybrid format that supports focus and collaboration.

Compensation & Benefits

Competitive, market-based pay, benchmarked by role and location — plus health coverage, paid time off, wellness support, and learning opportunities.

People-first Leadership

Approachable leaders who communicate openly, keep teams close to the strategy, and support long-term planning.

Advanced tech communities

Stay close to AI/ML, Cloud, Quantum Computing, IoT, and Robotics communities, with projects built on modern frameworks.

More opportunities available

Browse all open positions to find the best fit for your experience.

Browse all positions
101189