⚡ New

Inference & Performance Engineer @ Innowise

Innowise

WarszawaFull-timeMid LevelOn-site

Job Description


Core (must-have):

  • Strong Python and solid general software engineering fundamentals
  • Hands-on experience deploying at least one ML/LLM model to production inference — cloud serving or edge, either counts
  • Working knowledge of at least one inference/serving framework (vLLM, Triton, TensorRT-LLM, ONNX Runtime, llama.cpp/ggml, TGI, or similar)
  • Practical understanding of core optimization techniques: quantization, batching, caching, graph- or kernel-level optimization
  • Comfortable reasoning about latency/throughput/memory trade-offs
  • Solid grasp of deep learning fundamentals and transformer architectures

Strong plus — this is your specialization axis, not a day-one requirement:
  • Production C++ experience, especially for edge/on-device or runtime-level work
  • CUDA / GPU kernel programming exposure
  • Direct experience with llama.cpp, ggml, TensorRT-LLM, SGLang, FlashInfer, or similar low-level inference engines
  • Kubernetes / cloud infrastructure experience for GPU workloads
  • Experience with diffusion models
,(Deploy and optimize ML/LLM models for production inference across cloud GPU and, on select projects, edge/on-device targets, Work with inference/serving frameworks — vLLM, Triton Inference Server, TensorRT-LLM, ONNX Runtime, or llama.cpp/ggml, depending on the project's stack, Apply optimization techniques: quantization, pruning/distillation, operator fusion, graph/kernel-level compilation, KV-cache and batching strategies, Profile and tune runtime performance — latency, throughput, memory footprint, startup time, stability under long-running sessions, Build and maintain inference infrastructure: containerized deployment, GPU scheduling (Kubernetes), autoscaling, observability, benchmarking pipelines, On select engagements: work directly in C++ inference runtimes (e.g., llama.cpp/ggml-style engines), including custom CUDA kernel work, for edge and on-device deployment, Partner with research/ML engineers to take models from prototype to production, and with client engineering teams on integration) Requirements: Python, ML, LLM, C++, CUDA, GPU, Kubernetes

Posted Today

Related Jobs

Related Searches

Apply Now