(APIServer pid=5990) INFO 09-27 09:45:11 [api_utils.py:347] (APIServer pid=5990) INFO 09-27 09:45:11 [api_utils.py:347] █ █ █▄ ▄█ (APIServer pid=5990) INFO 09-27 09:45:11 [api_utils.py:347] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.30.0 (APIServer pid=5990) INFO 09-27 09:45:11 [api_utils.py:347] █▄█▀ █ █ █ █ model /workspace/models/awq (APIServer pid=5990) INFO 09-27 09:45:11 [api_utils.py:347] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀ (APIServer pid=5990) INFO 09-27 09:45:11 [api_utils.py:347] (APIServer pid=5990) INFO 09-27 09:45:11 [api_utils.py:286] non-default args: {'model_tag': '/workspace/models/awq', 'model': '/workspace/models/awq', 'max_model_len': 32768, 'served_model_name': ['q'], 'attention_backend': 'FLASH_ATTN', 'limit_mm_per_prompt': {'image': 0, 'video': 0}, 'max_num_seqs': 64} (APIServer pid=5990) INFO 09-27 09:45:11 [model.py:692] Resolved architecture: Qwen3_5ForConditionalGeneration (APIServer pid=5990) INFO 09-27 09:45:11 [model.py:2030] Using max model len 32768 (APIServer pid=5990) INFO 09-27 09:45:12 [registry.py:146] All limits of multimodal modalities supported by the model are set to 0, running in text-only mode. (APIServer pid=5990) INFO 09-27 09:45:12 [model.py:900] Disabled mm_prefix attention mode because multimodal inputs are configuration-disabled. Attention backends without mm_prefix support may now be selected. (APIServer pid=5990) INFO 09-27 09:45:12 [config.py:625] Mamba cache mode is set to 'align' for Qwen3_5ForConditionalGeneration by default when prefix caching is enabled (APIServer pid=5990) INFO 09-27 09:45:12 [kernel.py:416] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'], gelu_and_mul_sparse=['triton', 'native']) (EngineCore pid=6120) INFO 09-27 09:45:17 [core.py:123] Initializing a V1 LLM engine (v0.30.0) with config: model='/workspace/models/awq', speculative_config=None, tokenizer='/workspace/models/awq', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=compressed-tensors, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=q, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none'], 'ir_enable_torch_wrap': True, 'splitting_ops': ['vllm::unified_attention_with_output', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::qwen4_exp_ple_short_conv', 'vllm::qwen4_exp_qsa_with_output', 'vllm::linear_attention', 'vllm::qwen_gdn_attention_core', 'vllm::qwen_gdn_attention_core_fused_norm_packed', 'vllm::gdn_attention_core_xpu', 'vllm::olmo_hybrid_gdn_full_forward', 'vllm::sparse_attn_indexer', 'vllm::rocm_aiter_sparse_attn_indexer', 'vllm::deepseek_v4_attention', 'vllm::hpc_rope_norm_forward', 'vllm::unified_kv_cache_update', 'vllm::unified_mla_kv_cache_update'], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2, 4, 8, 16, 24, 32, 40, 48, 56, 64, 72, 80, 88, 96, 104, 112, 120, 128], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': 128, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'], gelu_and_mul_sparse=['triton', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, moe_backend='auto', sparse_indexer_topk_backend='auto', linear_backend='auto', linear_backend_per_quant=None) (EngineCore pid=6120) INFO 09-27 09:45:17 [registry.py:146] All limits of multimodal modalities supported by the model are set to 0, running in text-only mode. (EngineCore pid=6120) INFO 09-27 09:45:17 [parallel_state.py:1827] world_size=1 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_2649342d4a0f4c68bba7af713b9a3dd9 backend=nccl (EngineCore pid=6120) INFO 09-27 09:45:17 [parallel_state.py:2267] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, ETP rank 0, EP rank N/A, EPLB rank N/A (EngineCore pid=6120) INFO 09-27 09:45:17 [gpu_worker.py:441] Using V2 Model Runner (EngineCore pid=6120) INFO 09-27 09:45:18 [model_runner.py:396] Loading model from scratch... (EngineCore pid=6120) INFO 09-27 09:45:18 [cuda.py:597] Using backend AttentionBackendEnum.FLASH_ATTN for vit attention (EngineCore pid=6120) INFO 09-27 09:45:18 [mm_encoder_attention.py:372] Using AttentionBackendEnum.FLASH_ATTN for MMEncoderAttention. (EngineCore pid=6120) INFO 09-27 09:45:18 [compressed_tensors_wNa16.py:132] Using MarlinLinearKernel for CompressedTensorsWNA16 (EngineCore pid=6120) INFO 09-27 09:45:18 [qwen_gdn_linear_attn.py:176] Using FlashInfer GDN prefill kernel (requested=auto, head_k_dim=128). (EngineCore pid=6120) INFO 09-27 09:45:18 [qwen_gdn_linear_attn.py:528] GDN decode kernel: cuda (EngineCore pid=6120) INFO 09-27 09:45:18 [cuda.py:478] Using AttentionBackendEnum.FLASH_ATTN backend. (EngineCore pid=6120) INFO 09-27 09:45:18 [flash_attn.py:1116] Using FlashAttention version 2 (EngineCore pid=6120) Failed to get device capability: SM 12.x requires CUDA >= 12.9. (EngineCore pid=6120) Failed to get device capability: SM 12.x requires CUDA >= 12.9. (EngineCore pid=6120) INFO 09-27 09:45:18 [weight_utils.py:895] Filesystem type for checkpoints: OVERLAY. Checkpoint size: 19.57 GiB. Available RAM: 20.53 GiB. (EngineCore pid=6120) INFO 09-27 09:45:18 [weight_utils.py:925] Auto-prefetch is disabled because the filesystem (OVERLAY) is not a recognized network FS (NFS/Lustre) and the checkpoint size (19.57 GiB) exceeds 90% of available RAM (20.53 GiB). (EngineCore pid=6120) Loading safetensors checkpoint shards: 0% Completed | 0/5 [00:00= mamba page size. (EngineCore pid=6120) INFO 09-27 09:45:23 [interface.py:957] Padding mamba page size by 0.13% to ensure that mamba page size and attention page size are exactly equal. (EngineCore pid=6120) INFO 09-27 09:45:23 [utils.py:320] Using LBNHC KV cache layout. (EngineCore pid=6120) INFO 09-27 09:45:29 [backends.py:1094] Using cache directory: /root/.cache/vllm/torch_compile_cache/7bc2ac9422/rank_0_0/backbone for vLLM's torch.compile (EngineCore pid=6120) INFO 09-27 09:45:29 [backends.py:1155] Dynamo bytecode transform time: 5.60 s (EngineCore pid=6120) INFO 09-27 09:45:38 [backends.py:393] Compiling a graph for compile range (1, 2048) takes 9.33 s (EngineCore pid=6120) INFO 09-27 09:45:41 [backends.py:920] collected artifacts: 65 entries, 21 artifacts, 27467840 bytes total (EngineCore pid=6120) INFO 09-27 09:45:41 [decorators.py:719] saved AOT compiled function to /root/.cache/vllm/torch_compile_cache/torch_aot_compile/02cb00ab55408bc302d890c919296410ad44d2b6f096bf98beb9ee9d39c3b323/rank_0_0/model (EngineCore pid=6120) INFO 09-27 09:45:41 [monitor.py:53] torch.compile took 17.81 s in total (EngineCore pid=6120) INFO 09-27 09:45:47 [monitor.py:81] Initial profiling/warmup run took 5.96 s (EngineCore pid=6120) Capturing CUDA graphs (PIECEWISE): 0%| | 0/19 [00:00