Inference Endpoints (dedicated) documentation

Pricing

Hugging Face's logo
Join the Hugging Face community

and get access to the augmented documentation experience

to get started

Pricing

When you create an Endpoint, you can select the instance type to deploy and scale your model according to an hourly rate. Inference Endpoints is accessible to Hugging Face accounts with an active subscription and credits added. We recommend enabling automatic recharge to avoid service disruptions after credits are exhausted. A user or organization account will be charged for the compute resources used while successfully deployed Endpoints (ready to serve) are initializing and in a running state.

Below, you can find the hourly pricing for all available instances and accelerators, and examples of how costs are calculated: While the prices are shown by the hour, the actual cost is billed per minute.

CPU Instances

The table below shows currently available CPU instances and their hourly pricing. If the instance type cannot be selected in the application, you need to request quota to use it.

ProviderInstance TypeInstance SizeHourly ratevCPUsMemoryArchitecture
awsintel-sprx1$0.03312 GBIntel Sapphire Rapids
awsintel-sprx2$0.06724 GBIntel Sapphire Rapids
awsintel-sprx4$0.13448 GBIntel Sapphire Rapids
awsintel-sprx8$0.268816 GBIntel Sapphire Rapids
awsintel-sprx16$0.5361632 GBIntel Sapphire Rapids
azureintel-xeonx1$0.06012 GBIntel Xeon
azureintel-xeonx2$0.12024 GBIntel Xeon
azureintel-xeonx4$0.24048 GBIntel Xeon
azureintel-xeonx8$0.480816 GBIntel Xeon
gcpintel-sprx1$0.05012 GBIntel Sapphire Rapids
gcpintel-sprx2$0.10024 GBIntel Sapphire Rapids
gcpintel-sprx4$0.20048 GBIntel Sapphire Rapids
gcpintel-sprx8$0.400816 GBIntel Sapphire Rapids
awsintel-iclx1$0.03212 GBIntel Ice Lake - Deprecated from July 2025
awsintel-iclx2$0.06424 GBIntel Ice Lake - Deprecated from July 2025
awsintel-iclx4$0.12848 GBIntel Ice Lake - Deprecated from July 2025
awsintel-iclx8$0.256816 GBIntel Ice Lake - Deprecated from July 2025

GPU Instances

The table below shows currently available GPU instances and their hourly pricing. If the instance type cannot be selected in the application, you need to request quota to use it.

ProviderInstance TypeInstance SizeHourly rateGPUsMemoryArchitecture
awsnvidia-t4x1$0.5114 GBNVIDIA T4
awsnvidia-t4x4$3456 GBNVIDIA T4
awsnvidia-l4x1$0.8124 GBNVIDIA L4
awsnvidia-l4x4$3.8496 GBNVIDIA L4
awsnvidia-a10gx1$1124 GBNVIDIA A10G
awsnvidia-a10gx4$5496 GBNVIDIA A10G
awsnvidia-l40sx1$1.8148 GBNVIDIA L40S
awsnvidia-l40sx4$8.34192 GBNVIDIA L40S
awsnvidia-l40sx8$23.58384 GBNVIDIA L40S
awsnvidia-a100x1$2.5180 GBNVIDIA A100
awsnvidia-a100x2$52160 GBNVIDIA A100
awsnvidia-a100x4$104320 GBNVIDIA A100
awsnvidia-a100x8$208640 GBNVIDIA A100
awsnvidia-h100x1$4.5180 GBNVIDIA H100 - Deprecated from December 2025
awsnvidia-h100x2$92160 GBNVIDIA H100 - Deprecated from December 2025
awsnvidia-h100x4$184320 GBNVIDIA H100 - Deprecated from December 2025
awsnvidia-h100x8$368640 GBNVIDIA H100 - Deprecated from December 2025
awsnvidia-h200x1$51141 GBNVIDIA H200
awsnvidia-h200x2$102282 GBNVIDIA H200
awsnvidia-h200x4$204564 GBNVIDIA H200
awsnvidia-h200x8$4081128 GBNVIDIA H200
awsnvidia-b200x1$9.251256 GBNVIDIA B200 - Deprecated from December 2025
awsnvidia-b200x2$18.52512 GBNVIDIA B200 - Deprecated from December 2025
awsnvidia-b200x4$3741024 GBNVIDIA B200 - Deprecated from December 2025
awsnvidia-b200x8$7482048 GBNVIDIA B200 - Deprecated from December 2025
gcpnvidia-t4x1$0.5116 GBNVIDIA T4
gcpnvidia-l4x1$0.7124 GBNVIDIA L4
gcpnvidia-l4x4$3.8496 GBNVIDIA L4
gcpnvidia-a100x1$3.6180 GBNVIDIA A100
gcpnvidia-a100x2$7.22160 GBNVIDIA A100
gcpnvidia-a100x4$14.44320 GBNVIDIA A100
gcpnvidia-a100x8$28.88640 GBNVIDIA A100
gcpnvidia-h100x1$10180 GBNVIDIA H100
gcpnvidia-h100x2$202160 GBNVIDIA H100
gcpnvidia-h100x4$404320 GBNVIDIA H100
gcpnvidia-h100x8$808640 GBNVIDIA H100

INF2 Instances

The table below shows currently available INF2 instances and their hourly pricing. If the instance type cannot be selected in the application, you need to request quota to use it.

ProviderInstance TypeInstance SizeHourly rateAcceleratorsAccelerator MemoryRAMArchitecture
awsinf2x1$0.75132 GB14.5 GBAWS Inferentia2
awsinf2x12$1212384 GB760 GBAWS Inferentia2

Pricing examples

The following example pricing scenarios demonstrate how costs are calculated. You can find the hourly rate for all instance types and sizes in the tables above. Use the following formula to calculate the costs:

instance hourly rate * ((hours * # min replica) + (scale-up hrs * # additional replicas))

Basic Example

  • AWS CPU intel-spr x2 (2x vCPUs 4GB RAM)
  • Autoscaling (minimum 1 replica, maximum 1 replica)

hourly cost

instance hourly rate * (hours * # min replica) = hourly cost
$0.067/hr * (1hr * 1 replica) = $0.067/hr

monthly cost

instance hourly rate * (hours * # min replica) = monthly cost
$0.064/hr * (730hr * 1 replica) = $46.72/month

basic-chart

Advanced Example

  • AWS GPU small (1x GPU 14GB RAM)
  • Autoscaling (minimum 1 replica, maximum 3 replica), every hour a spike in traffic scales the Endpoint from 1 to 3 replicas for 15 minutes

hourly cost

instance hourly rate * ((hours * # min replica) + (scale-up hrs * # additional replicas)) = hourly cost
$0.5/hr * ((1hr * 1 replica) + (0.25hr * 2 replicas)) = $0.75/hr

monthly cost

instance hourly rate * ((hours * # min replica) + (scale-up hrs * # additional replicas)) = monthly cost
$0.5/hr * ((730hr * 1 replica) + (182.5hr * 2 replicas)) = $547.5/month

advanced-chart

Update on GitHub