⚡ New
Inference & Performance Engineer @ Innowise
Innowise
WarszawaFull-timeMid LevelOn-site
Job Description
Core (must-have):
- Strong Python and solid general software engineering fundamentals
- Hands-on experience deploying at least one ML/LLM model to production inference — cloud serving or edge, either counts
- Working knowledge of at least one inference/serving framework (vLLM, Triton, TensorRT-LLM, ONNX Runtime, llama.cpp/ggml, TGI, or similar)
- Practical understanding of core optimization techniques: quantization, batching, caching, graph- or kernel-level optimization
- Comfortable reasoning about latency/throughput/memory trade-offs
- Solid grasp of deep learning fundamentals and transformer architectures
Strong plus — this is your specialization axis, not a day-one requirement:
- Production C++ experience, especially for edge/on-device or runtime-level work
- CUDA / GPU kernel programming exposure
- Direct experience with llama.cpp, ggml, TensorRT-LLM, SGLang, FlashInfer, or similar low-level inference engines
- Kubernetes / cloud infrastructure experience for GPU workloads
- Experience with diffusion models
Posted Today