⚡ Spark-X2.5-4B

Chat with Spark-X2.5-4B — a compact general-purpose LLM from XHToken with hybrid sliding-window/full attention, a native 1M-token context window, 200+ languages, and strong reasoning, coding, and agent skills.

Thinking mode (on by default) lets the model reason in a collapsible 💭 block before answering — better on hard problems, slower on chit-chat. Runs in bf16 on ZeroGPU; this demo caps input at 16K tokens. Each request uses your ZeroGPU allowance — if you hit a quota or duration error, sign in to Hugging Face or lower Max new tokens.