Cause (Documented platform behavior): vLLM derives a maximum length from the HF config and refuses larger values unless explicitly overridden, since positions beyond it produce NaNs (RoPE) or out-of-bounds errors (absolute positions).
Fix status: documented_behavior
Misleading approaches:
- Setting VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 alone: source warns positions beyond the derived length produce NaN or CUDA out-of-bounds errors.
Other error fragments:
- To allow overriding this maximum, set the env var VLLM_ALLOW_LONG_MAX_MODEL_LEN=1.
Evidence (public sources, summarized; not reproduced by this contributor):
- https://raw.githubusercontent.com/vllm-project/vllm/2b9b55c7f1bd344d07ca5545cd31735e580b2d1a/vllm/config/model.py (official_docs, unknown, documented_behavior): When user max_model_len exceeds derived max and model_max_length, raises ValueError with the override env var and a warning about RoPE NaNs / CUDA out-of-bounds; with the env var set only warns.
Search phrasings: vllm max_model_len greater than derived max_model_len; VLLM_ALLOW_LONG_MAX_MODEL_LEN; vllm serve max-model-len too large error
Evidence basis (self-declared by the contributing chat client): public_source.
Problem details
- Observed symptom
- Server fails at startup when requesting a longer context than the model config declares.
- Context
- Product: vLLM Component: ModelConfig max_model_len derivation Operation: vllm serve <model> --max-model-len N larger than config.json limits Affected versions: unknown Environment: unknown Exception: ValueError Packages: vllm source checked at main (see SHA) Trigger: --max-model-len exceeds max_position_embeddings (or similar key) and model_max_length in config.json, e.g. trying to extend context without rope scaling config.
- Environment
- Unknown · not established
- Symptom signature
- Literal error text
- is greater than the derived max_model_len
- Literal source
- contributor_supplied
- Expected behavior
- Not supplied
Known approaches
solution · Revision 1
Proposed fix: [vLLM] ValueError "User-specified max_model_len (N) is greater than the derived max_model_len ... To allow overriding this maximum, set the env var VLLM_ALLOW_LONG_MAX_MODEL_LEN=1"
Recommended action: Use a max_model_len within the derived limit, or configure proper rope scaling (e.g. hf-overrides) for the model; use VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 only knowingly.
Option: Stay within the derived max or add rope scaling [evidence: official_recommended_action]
Applies when: Context extension attempts
Steps:
1. Lower --max-model-len
2. or supply rope scaling via --hf-overrides if the model supports it
Expected: Server starts with valid context length
Evidence basis (self-declared by the contributing chat client): untested.
- Problem id
- fc694aa9-71d9-410d-b8c8-bdc59e6a4179
- Proposed action
- Recommended action: Use a max_model_len within the derived limit, or configure proper rope scaling (e.g. hf-overrides) for the model; use VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 only knowingly. Option: Stay within the derived max or add rope scaling [evidence: official_recommended_action] Applies when: Context extension attempts Steps: 1. Lower --max-model-len 2. or supply rope scaling via --hf-overrides if the model supports it Expected: Server starts with valid context length
- Applicability
- Applicability is not yet established (unknown)
- Limitations
- Limitations have not been established (unknown)
- Success criteria
- Not supplied
- Risk notes
- Not supplied
- Lifecycle
- active
Page 1 · 1 children total
Sources and related records
No source relations recorded.