Replicate

Replicate pricing charges per second of compute, and the per-second rate depends on which GPU your model runs on.

Pricing Model:

Pricing Model:

Pure usage, prepaid

Pure usage, prepaid

Pure usage, prepaid

Packaging Model:

Packaging Model:

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Credit Model:

Credit Model:

Prepaid, valid 1 year, non-refundable, auto reload available

Prepaid, valid 1 year, non-refundable, auto reload available

Prepaid, valid 1 year, non-refundable, auto reload available

Updated on:

Replicate pricing: paying by the second, until the model is official

Replicate pricing charges per second of compute, and the per-second rate depends on which GPU your model runs on. An A100 costs $0.001400 a second, an H100 $0.001525. But a subset of models, the ones Replicate calls official, ignore the clock entirely and bill per output image, per token or per second of generated video. Which meter applies isn't a setting you choose. It's a property of the model you picked.

Key takeaways

  • Replicate's single-GPU rates run from $0.000225 a second on a T4 to $0.001525 a second on an H100, with an 8x H100 node at $0.012200, all published on the pricing page with no login.

  • Official models bill per output unit instead of per second: $0.04 per image on FLUX 1.1 Pro, $0.09 per second of output video on Wan 2.1 i2v 480p.

  • On public models you pay only for active processing time. Setup and idle time are free, so cold boots cost you latency but not money.

  • On private models and deployments you pay for setup, idle and active time, because the hardware is dedicated to you.

Replicate pricing in 2026

Hardware

Per second

Per hour

Spec

 

CPU (small)

$0.000025

$0.09

1x CPU, 2GB RAM

Nvidia T4

$0.000225

$0.81

16GB GPU RAM

Nvidia L40S

$0.000975

$3.51

48GB GPU RAM

Nvidia A100 80GB

$0.001400

$5.04

80GB GPU RAM

Nvidia H100

$0.001525

$5.49

80GB GPU RAM

Nvidia H200

$0.001525

$5.49

Committed spend contract only

8x Nvidia H100

$0.012200

$43.92

Committed spend contract only

What Replicate actually meters

Replicate runs two meters, and the model decides which one you get.

Most community and private models meter wall-clock seconds of compute at the hardware's published rate. Every run creates a prediction, and Replicate charges the seconds that prediction spends actually executing.

Official models replace that with output units. FLUX 1.1 Pro charges $0.04 per output image, FLUX Schnell $3.00 per thousand output images, DeepSeek R1 $3.75 per million input tokens and $0.01 per thousand output tokens, and Wan 2.1 i2v 480p $0.09 per second of output video. Replicate named this category on 29 January 2025 and described the switch plainly: instead of being charged for the time a model runs, you're charged by output.

Two wrinkles matter. Models that call other models bill you for the root model's compute plus every downstream model it invokes, and the downstream list appears in the model's pricing section. And when Nano Banana Pro falls back to Seedream 5.0 lite under rate limiting, you pay the fallback model's price, not Nano Banana Pro's.

How credits work

Since 16 July 2025, every new Replicate account bills through prepaid credit. You buy a balance, and usage deducts from it.

Purchased credit is valid for one year from the purchase date and isn't refundable. Replicate documents no rollover concept, because the balance is simply money you've already spent.

Auto reload tops the balance back up when it drops past a threshold you set. The minimum threshold is $5 and the minimum reload balance is $15, and Replicate's own example is a $10 threshold with a $50 reload, which adds $40 when you hit $10. If your balance already sits at or below the threshold when you save, the reload fires immediately.

Accounts created before 16 July 2025 can stay on monthly arrears billing, where Replicate charges the previous month's usage at the start of the next one. Replicate says it intends to migrate most accounts to prepaid eventually. Credit is account-scoped, and organizations get a shared balance rather than pooling individual ones.

What happens when you hit the limit

Replicate throttles you before it stops you. As your credit balance approaches zero, it applies progressively stronger rate limits so you have time to top up rather than getting cut off without warning. The docs recommend keeping the balance above $20 via auto reload.

At zero, Replicate prevents new work from starting and shuts down any infrastructure running for you. A prediction can occasionally run past the balance, and Replicate charges the outstanding amount to your default payment method at month end.

Normal rate limits sit at 600 prediction creations a minute and 3,000 requests a minute elsewhere, with short bursts allowed. Accounts holding granted credit with no payment method on file get 1 request per second and 6 a minute. Predictions time out at 30 minutes.

How Replicate's pricing has changed across all these years

Date

Milestone

Source

 

2 Mar 2026

Nano Banana Pro fallback bills at the fallback model's rate

Vendor

21 Nov 2025

Approximate per-run cost shown on prediction and training pages

Vendor

8 Oct 2025

Invoice PDFs available from billing settings via Stripe

Vendor

26 Sep 2025

Low-balance throttling documented as a spend guardrail

Vendor

16 Jul 2025

All new accounts moved to prepaid credit. Existing accounts keep monthly arrears

Vendor

29 Jan 2025

Official models named and switched from time-based to per-output billing

Vendor

22 Nov 2024

Per-video pricing support added, alongside L40S GPUs and preview hardware pricing

Vendor

Flexprice’s Take

Replicate publishes one of the cleanest inference rate cards in this index, and the free-setup rule on public models is the detail worth copying.

Replicate doesn't charge you for cold boots, and that's rarer than it should be. On public models you pay only for seconds a prediction spends actively processing, so weight loading costs you latency and nothing else. Private models and deployments flip the rule, and the docs say so plainly instead of burying it.

The split meter is the hard part. One model in your workflow bills GPU-seconds at $0.001525 on an H100, the next bills $0.04 per output image on FLUX 1.1 Pro, and nothing in the API response tells you which axis you're on.

Forecast by model identity, not by volume. Models that call other models compound it, because your invoice includes compute you never invoked.

Best For

Teams running bursty, varied workloads across many models, where per-second billing beats a reserved GPU.

Watch Out For

Deployments left with a minimum instance count above zero, which bill idle time at the full hardware rate.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling inference and need per-model cost and margin per customer?

Flexprice meters it.

Flexprice’s Take

Replicate publishes one of the cleanest inference rate cards in this index, and the free-setup rule on public models is the detail worth copying.

Replicate doesn't charge you for cold boots, and that's rarer than it should be. On public models you pay only for seconds a prediction spends actively processing, so weight loading costs you latency and nothing else. Private models and deployments flip the rule, and the docs say so plainly instead of burying it.

The split meter is the hard part. One model in your workflow bills GPU-seconds at $0.001525 on an H100, the next bills $0.04 per output image on FLUX 1.1 Pro, and nothing in the API response tells you which axis you're on.

Forecast by model identity, not by volume. Models that call other models compound it, because your invoice includes compute you never invoked.

Best For

Teams running bursty, varied workloads across many models, where per-second billing beats a reserved GPU.

Watch Out For

Deployments left with a minimum instance count above zero, which bill idle time at the full hardware rate.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling inference and need per-model cost and margin per customer?

Flexprice meters it.

Customer
Sentiment Highlights

"On replicate.com a single image takes 1.5s at a price of 1000 images per $1."

aleyan, Hacker News, December 2025

"Neither fal nor replicate return accurate pricing in the response body."

seblavoie, Hacker News, February 2026

Frequently Asked Questions

Frequently Asked Questions

How much does Replicate cost per hour?

Does Replicate charge for cold boots?

Do Replicate credits expire?

What happens if you run out of credit on Replicate?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 400+ Builders on Slack

Join the Flexprice Community on Slack