Context-Dependent Effects of Dynamic Confidence Thresholds for Speculative Decoding on a Consumer GPU
Abstract
This technical note reports a consumer-GPU case study of generated-string-aware confidence thresholds for DSpark speculative decoding. On a six-source long-context holdout using Qwen3-8B Q4_K_M and an AMD Radeon RX 6800, the dynamic policy reached 1.0224x fixed-DSpark decode throughput while reducing proposed draft tokens by 12.3%. On 164 HumanEval Python tasks, fixed and dynamic DSpark passed the same 130 tasks, while the dynamic policy was slightly slower on short-function timing. The claims are limited to the tested model, runtime, device, and workloads. Generative AI assistance is fully disclosed in the manuscript. This work was conducted independently and was not sponsored, supervised, or endorsed by Green Hope High School.
// Source
Authors: Vicente Bertolotti