Inference & Performance Engineer
Core (must-have):
Strong plus — this is your specialization axis, not a day-one requirement:
- Strong Python and solid general software engineering fundamentals
- Hands-on experience deploying at least one ML/LLM model to production inference — cloud serving or edge, either counts
- Working knowledge of at least one inference/serving framework (vLLM, Triton, TensorRT-LLM, ONNX Runtime, llama.cpp/ggml, TGI, or similar)
- Practical understanding of core optimization techniques: quantization, batching, caching, graph- or kernel-level optimization
- Comfortable reasoning about latency/throughput/memory trade-offs
- Solid grasp of deep learning fundamentals and transformer architectures
Strong plus — this is your specialization axis, not a day-one requirement:
- Production C++ experience, especially for edge/on-device or runtime-level work
- CUDA / GPU kernel programming exposure
- Direct experience with llama.cpp, ggml, TensorRT-LLM, SGLang, FlashInfer, or similar low-level inference engines
- Kubernetes / cloud infrastructure experience for GPU workloads
- Experience with diffusion models
Podobne oferty
Site Reliability Engineer
Link Group
WARSZAWA
2026-09-22
DevOps Site Reliability Engineer
Link Group
WARSZAWA
2026-09-22
Site Reliability Engineer
Mindbox Sp. z o.o.
WARSZAWA
2026-09-22
Site Reliability Engineer SRE
Citi
WARSZAWA
2026-09-22
DevOps Engineer Azure
Innowise
WARSZAWA WARSAW
2026-09-22
Cloud Engineer Fully Remote
Mercor
WARSAW
2026-09-21
Junior DevOps Engineer
Klient TeamQuest
WARSZAWA
2026-09-21
Cloud Engineer Fully Remote
Mercor
WARSAW
2026-09-21
DevOps Engineer
Scalo
WARSZAWA
2026-09-25
Azure Platform Engineer
Mindbox Sp. z o.o.
WARSZAWA
2026-09-24