Follow us
Breaking
Tech News

Cloudflare AI Workers pricing shifts

Cloudflare has replaced neuron units with task-based pricing for LLMs, images, and audio. New AI Gateway tiers offer significant savings, making Cloudflare up to 56x cheaper than OpenAI for certain workloads at the $200 tier.

Share

Goodbye to neurons

Cloudflare replaced the neuron unit with pricing based on specific model tasks. This transition removes the difficulty users faced when comparing different model types. LLM costs depend on input and output tokens. Image generation models charge based on output resolution and the number of steps. Audio models use seconds of audio input for billing. This change follows feedback that neurons were difficult to grasp.

Cloudflare removed neurons.

The free tier provides 10,000 tokens a day for text, 250 steps for image generation up to 1024×1024 resolution, and 10 minutes of audio per day. The model catalog now splits into a static catalog and a dynamic catalog. Static models remain curated by Cloudflare to ensure availability and speed. Dynamic models come from the "Run Any Model" feature. Llama 3.2 1B costs $0.027 per million input tokens and $0.201 per million output tokens. Llama 3.3 70B costs $0.293 per million input tokens and $2.253 per million output tokens. Granite 4.0 Micro costs $0.017 per million input tokens and $0.112 per million output tokens. Deepseek V4 Flash costs $0.440 per million input tokens and $1.320 per million output tokens.

Task Type Pricing Metric
Text (LLM) Input and output tokens
Image Generation Resolution and steps
Speech-to-text Seconds of audio input
Embeddings Input tokens

AI Gateway request tiers

Cloudflare AI Gateway manages traffic between applications and foundation model providers. The service allows users to pay for third-party model usage through a Cloudflare invoice. This feature includes a small transaction convenience fee. The 2026 pricing model uses request-based tiers instead of per-token fees. This structure aligns incentives with concise prompts. The gateway supports providers like OpenAI, Anthropic, Google AI Studio, Hugging Face, and Replicate.

The 2026 tiers scale by monthly request allowances.

Plan Tier Monthly Request Allowance Cost per Request
Free 10,000 requests daily $0
$25 100,000 requests $0.00025
$50 500,000 requests $0.0001
$100 2,000,000 requests $0.00005
$200 5,000,000 requests $0.00004
$300 10,000,000 requests $0.00003

Log retention impacts costs. The free tier includes 100,000 logs per month across all gateways. Paid plans increase this to 1,000,000 logs. If you need to store more logs, you pay $8 per 100,000 logs per month. Logpush streaming to external storage costs $0.05 per million records after the first 10 million records on paid plans. The gateway also enables request routing and rate limiting to control traffic flow.

You should watch your usage.

Edge performance and costs

The edge deployment reduces latency for global users. Cloudflare runs inference near the user. This proximity cuts latency compared to centralized APIs. The edge deployment covers over 300 cities globally. Users in North America see 100ms to 200ms latency.

Cloudflare remains the cheapest inference option in 2026.

A typical request for OpenAI GPT-4o with 500 input tokens and 100 output tokens costs $0.00225; this makes Cloudflare 9x cheaper than OpenAI for this specific workload. Cloudflare’s $25 tier for 100,000 requests costs only $0.00025 per request. This makes Cloudflare 9x cheaper than OpenAI. At the $200 tier, Cloudflare costs 56x less than OpenAI. Anthropic Sonnet 4.6 costs $0.003 for the same 500 input and 100 output tokens. This makes Anthropic 12x more expensive than Cloudflare.

Provider Metric Price
OpenAI GPT-4o 1M Input Tokens $2.50
OpenAI GPT-4o 1M Output Tokens $10.00
Cloudflare ($25 Tier) 100k Requests $25.00

European users see similar ranges. OpenAI latency is 800ms to 1200ms. Anthropic latency is 700ms to 1000ms. Cloudflare latency for US East is 180ms to 350ms. Vectorize v2 also improved performance, reducing median latency from 500ms to 30ms. The lack of fine-grained control and fixed usage allocations can hinder budgeting and compliance. Many organizations find that they cannot certify what happens inside the gateway due to the black-box routing.

Can teams handle the loss of visibility when they disable logging for privacy?

Share

Technewsdaily

Senior tech writer covering AI, gadgets and cybersecurity. Breaking down the news that matters, every day.