GPU inference compute without the cloud lock-in
Explore provider-run GPUs for LLM inference, inspect the details that matter, and compare available capacity before you begin a marketplace connection.
Looking to contribute hardware instead? Learn about earning from idle compute.
These counters come from the marketplace and can change as providers and jobs change.
Choose by workload
Find the right GPU inference capacity
Start with the model and hardware requirements of your workload, then compare individual listings rather than choosing an opaque bundle.
Available compute to explore
Live listing previews; availability and pricing can change.
From search to request
A clear path to your first inference call
- 01
Browse
Filter the marketplace by model, availability, provider, and instance type.
- 02
Inspect
Open a listing to review the provider, hardware, pricing, and connection details.
- 03
Connect
Use the marketplace connection flow and review the REST API resources relevant to your application.
- 04
Plan
Use the listing details and API reference to plan how your application will work with the marketplace.
What you can use today
- A marketplace for provider listings and model availability
- A marketplace for comparing live provider listings before you connect
- Displayed provider rates and a $0.10 per 1,000-token platform fee
- Public REST API documentation for listings, endpoints, connections, jobs, billing, and keys
GPU inference questions
The practical details to check before you connect an application.
Ready to compare GPU compute?
Browse the live marketplace or read the API reference before you connect an application.