Speculative decoding can accelerate LLM token generation by roughly 1.6x on structured tasks like coding and JSON output, but the speedup ...