You are viewing a preview of this job. Log in or register to view more details about this job.

Machine Learning Engineer

Make the models themselves faster to serve. You will work on speculative decoding, attention and KV cache optimizations, and model-level techniques that raise tokens per second while preserving full model quality.