Production AI inference
on wylon

A high-performance LLM inference platform built for developers and enterprises.

Core products · Token Factory

One API key for chat, image, and video

Generative AI inference for developers and enterprises. Three capabilities, one account, open for sign-up.

Chat and agents Live

Production chat inference for complex work

Built for long context, code, knowledge systems, and agent workflows, with streaming, tool calls, and system-level context caching.

GLM-5.3 DeepSeek-V4-Flash
API compatibilityOpenAI / Anthropic
BillingPer token
Image generation and editing Live

From prompt to usable visual

Generate marketing assets and rendered text across multiple aspect ratios, or edit images with complex instructions.

Qwen-Image-2512 SenseNova-U1.5-8B-MoT
API compatibilityOpenAI Images
BillingPer image
Video generation Live

Generate short-form video efficiently

For product demos, campaign content, and complex motion shots, with 480p and 720p output.

Wan2.2-T2V-A14B
Request formatVolcengine Ark API compatible
BillingPer second
Core products · GPU Supernode

GPU Supernode

GPU compute for online supernode rental and enterprise dedicated service. Now accepting early-access requests.

Two delivery options Coming soon

Online Supernode Rental

Online access to GPU Supernode resources with request-based provisioning, unified delivery, and room for future elastic expansion.

Enterprise Dedicated Service

Dedicated server-level GPU capacity with tailored delivery support, resource isolation, and deployment planning.

Core technology

Inference foundation for domestic GPUs

Token Factory and GPU Supernode share the same technology foundation. Through an LLM inference OS that coordinates vertical storage resources and horizontal GPU fabric, wylon turns domestic GPU capacity from isolated resources into system-level inference infrastructure.

Products & Platform

One foundation, two product forms

Token Factory delivers model inference directly through APIs;
GPU Supernode delivers compute resources for inference infrastructure.
Developers and enterprises can choose by usage pattern and resource needs, while both share the same system foundation.

wylon Supernode System architecture: horizontal GPU fabric and vertical storage plane converge at NoC, while Hiten LLM OS coordinates through software-hardware Co-Design

Extending the Pareto frontier of domestic GPUs

The wylon Supernode System improves single-user response speed while sustaining higher system throughput, extending the efficiency frontier for domestic GPUs in large-scale model inference scenarios.

Why wylon

Built for LLM inference workloads

Full-stack, hardware-software co-designed

End-to-end cloud infrastructure across GPU compatibility, super-node architecture, inference runtime, and API services - tightly integrated to reduce performance overhead.

Enterprise

Service quality

A high-availability architecture built for production workloads, delivering dependable, enterprise-grade service quality.

0

Ops overhead

Drop-in inference API - no infrastructure to manage.

10×

Speedup

Backed by system-level caching, prefill is up to 10× faster than baseline implementations.

6+

Chip families supported

wylon's super-node system supports Biren, MetaX, Sunrise, Cambricon, and more.

Learn more

Explore the platform, APIs, and deployment options.

Read the docs
Silicon partners
Sunrise Sunrise MetaX MetaX Biren Biren Cambricon Cambricon
Get started

Try the wylon Token Factory

Token Factory is open for sign-ups and can be integrated in minutes.