Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
nsagent
on April 21, 2024
|
parent
|
context
|
favorite
| on:
Lossless Acceleration of LLM via Adaptive N-Gram P...
How does this differ from the 2018 NeurIPS paper, Blockwise Parallel Decoding for Deep Autoregressive Models?
https://arxiv.org/abs/1811.03115
tripplyons
on April 22, 2024
[–]
They use a separate ngram model to generate the proposed sequence instead of extra heads on top of the main model. The process of verifying the proposed sequence appears to be the same.
Consider applying for YC's Fall 2026 batch!
Applications
are open till July 27.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
https://arxiv.org/abs/1811.03115