Make the models themselves faster to serve. You will work on speculative decoding, attention and KV cache optimizations, and model-level techniques that raise tokens per second while preserving full model quality.