Modal Serverless GPU: $30/Month Free Credits

Quick answer: Modal is a serverless GPU platform where you attach a GPU to a Python function with one decorator and are billed per second only while it runs — idle means $0. The free Starter plan hands you $30 in compute credits every month (recurring, not a one-time trial), enough for ~50 T4-hours or ~12 A100-80GB-hours. A small project that sits at zero between bursts can genuinely run free indefinitely.
✅ Verified 2026-08-31. We re-read Modal's own pricing page and the free-tier numbers below still match what it publishes:
Credit$30/month
Containers100
GPU concurrency10
Free tiers moved a lot in 2026 — several providers we cover quietly started requiring a credit card — so we re-check rather than assume.

You need a real datacenter GPU for a few minutes a day — to run a Llama model, fine-tune an embedding model, or generate a batch of images. Renting an instance by the hour means paying for the 23+ idle hours too. Modal deletes that tax: GPU count is zero when nothing calls it, and so is your bill.

Free Tier at a Glance

DimensionModal Starter (free)
Free compute$30 of credits every month, recurring; credits don't roll over
Credit card to start?No — needed only to spend beyond the credit
GPUs availableT4, L4, A10, L40S, A100 (40/80 GB), H100, H200, B200 — attach one per function
Idle cost$0 — containers scale to zero between requests
Sleeps / expires?Scales to zero when idle (cold start on wake); the $30 credit refreshes monthly
Commercial useAllowed on the Starter plan
RegionUS and EU datacenters; region pinning available in code — check current docs for specifics
What you deployAny Python function or container with a GPU attached — inference endpoints, batch jobs, fine-tuning, cron. No built-in web UI (bring your own or use HF Spaces for that)

What Is Modal?

Modal is a serverless cloud platform for compute-heavy, bursty, GPU-hungry code. Its defining idea: you describe infrastructure in the same Python file as your logic — no Dockerfile, no YAML, no separate ops layer. You declare the container image, GPU, secrets, and schedule as Python objects and decorators; Modal builds and runs it. Three properties make it fit the "I need a GPU occasionally" problem:

  • Per-second billing. An A100 used for 90 seconds costs 90 seconds of A100 time — not an hour, not a day.
  • Scale to zero (and back up). No requests means containers spin down to zero; the next request spins one back up. You never pay for idle, never have to remember to turn anything off.
  • Code-first infrastructure. modal run executes once; modal deploy turns it into a persistent service with a URL. That's the whole mental model.

It's used for serverless inference, massively parallel batch jobs, fine-tuning, cron tasks, and sandboxed code execution for agents. Full docs live in the Modal guide.

Is Modal Really Free?

Yes, with an honest asterisk: the free tier is a monthly credit, not unlimited service. The Starter plan costs $0, needs no upfront payment, and grants $30 of compute credits that refresh every month. Credits don't roll over; usage beyond $30 is billed at standard per-second rates.

PlanMonthly costIncluded creditsWho it's for
Starter$0$30 / monthIndividuals, side projects, evaluation
Team$250$100 / monthTeams needing more seats and higher limits
EnterpriseCustomCustomLarge orgs, compliance, dedicated support

"$30/month forever" is structurally different from a one-time signup credit. Replicate and RunPod are pay-as-you-go: after any small trial, every GPU-second is real money. Modal's $30 comes back on the 1st, so a well-behaved project that does a few GPU-hours a month and otherwise sits at zero can run at $0 indefinitely (per the Modal pricing page, June 2026).

What $30/Month Actually Buys You

Per-second rates converted to hourly, then to GPU-hours $30 buys each month (from the pricing page, June 2026):

GPUPrice / second≈ / hourFree hours on $30/moGood for
Nvidia T4 (16 GB)$0.000164~$0.59~50 hoursSmall models, embeddings, Whisper
Nvidia L4 (24 GB)$0.000222~$0.80~37 hours7B inference, image generation
Nvidia A10 (24 GB)$0.000306~$1.10~27 hours7B–13B inference, light fine-tuning
Nvidia L40S (48 GB)$0.000542~$1.95~15 hoursLarger image models, 13B–30B
Nvidia A100 40 GB$0.000583~$2.10~14 hours30B inference, serious fine-tuning
Nvidia A100 80 GB$0.000694~$2.50~12 hours70B inference (quantized), training
Nvidia H100$0.001097~$3.95~7.5 hoursFast 70B, heavy training
Nvidia H200$0.001261~$4.54~6.6 hoursLarge-context, big-model training
Nvidia B200$0.001736~$6.25~4.8 hoursFrontier-scale workloads

Two caveats. A GPU container also bills a little CPU ($0.0000131/core/s, min 0.125 cores) and memory ($0.00000222/GiB/s), so real GPU-hours run a hair below the table — but the GPU rate dominates. And "hours" means active hours: because Modal scales to zero, a web endpoint answering a few hundred requests a day might accumulate only 20–30 minutes of real GPU time, so ~50 T4-hours can cover a low-traffic service all month with room to spare. On the cheaper GPUs, $30/month is real working capital.

Your First Modal GPU Function in 5 Minutes

Install and authenticate (modal setup opens a browser to link — or create — your free account):

pip install modal
modal setup

Create hello_gpu.py — image, GPU, and logic all in one file:

import modal

image = modal.Image.debian_slim().pip_install("torch")
app = modal.App("hello-gpu", image=image)

@app.function(gpu="T4")          # attach a real Nvidia T4 to this function
def check_gpu():
    import torch
    name = torch.cuda.get_device_name(0)
    return f"Running on: {name}"

@app.local_entrypoint()
def main():
    print(check_gpu.remote())     # .remote() runs it in the cloud, not locally
modal run hello_gpu.py

Modal builds the image (cached afterward), provisions a T4, runs check_gpu there, streams back Running on: Tesla T4, and tears the container down. The GPU existed for a few seconds, you paid for a few seconds, nothing is left running to forget. The key line is check_gpu.remote(): .remote() ships the function to the cloud; calling it normally runs locally.

Deploy an LLM as a Web Endpoint That Scales to Zero

The other major pattern is a persistent HTTP endpoint backed by a GPU that spins up on the first request and scales back to zero when traffic stops. The idiomatic way uses a class so the model loads once per container and stays warm:

import modal

image = (
    modal.Image.debian_slim()
    .pip_install("transformers", "torch", "accelerate", "fastapi[standard]")
)
app = modal.App("free-llm-api", image=image)

@app.cls(gpu="L4")
class LLM:
    @modal.enter()                       # runs once per container, on startup
    def load(self):
        from transformers import pipeline
        self.pipe = pipeline(
            "text-generation",
            model="Qwen/Qwen2.5-1.5B-Instruct",
            device_map="auto",
        )

    @modal.fastapi_endpoint(method="POST")   # expose this method as an HTTP endpoint
    def generate(self, prompt: str):
        out = self.pipe(prompt, max_new_tokens=256)
        return {"output": out[0]["generated_text"]}
modal deploy llm_api.py

Modal prints a public URL. The @modal.enter() hook loads the 1.5B model into GPU memory once per container, and Modal keeps the container alive for a short idle window after the last request before scaling to zero — so a traffic burst reuses one warm container and a quiet night costs nothing. Call it like any API:

curl -X POST https://yourworkspace--free-llm-api-llm-generate.modal.run 
  -H "Content-Type: application/json" 
  -d '{"prompt": "Explain serverless GPUs in one sentence."}'

Need multiple routes or a full ASGI app? Modal also provides @modal.asgi_app, documented in the web endpoints guide.

What Modal Is Good At

  • Bursty serverless inference — an open model behind an endpoint that costs nothing when idle. Ideal for internal tools, demos, and low-to-medium-traffic apps.
  • Massively parallel batch jobs — use .map() to fan a function across hundreds of containers (embed a large corpus, transcribe an audio archive), then scale straight back to zero.
  • Fine-tuning and training — spin up an A100 or H100 for the length of a job, pay only those hours. No reserved instance, no commitment.
  • Scheduled cron tasks — add schedule=modal.Cron("0 8 * * *") and Modal runs it every morning with a GPU attached only when it fires.
  • Sandboxes for AI agents — a fresh isolated container to run untrusted code, a fit for tool-using agents.

Modal vs RunPod vs Replicate vs Hugging Face Spaces

PlatformRecurring free tierBillingInterfaceBest for
Modal$30/mo, every monthPer-second, scale to zeroPython code & decoratorsCustom AI backends, batch, training
RunPodNo recurring creditPer-second; pods or serverlessContainer templates & UIRaw GPU access, cheapest hourly
ReplicateNo recurring creditPer-second per predictionModel registry & APIRunning & sharing published models
HF Spaces (ZeroGPU)Free, time-sliced RTX Pro 6000 Blackwell (5 min/day)2 free ZeroGPU Spaces / PROGradio & Git pushPublic demos with a shareable URL

Replicate is easiest for calling a model someone else packaged, but you pay from the first prediction. RunPod gives the cheapest raw pods, but you manage more lifecycle and there's no monthly credit. HF Spaces is unbeatable for a free public Gradio demo with a URL, but not arbitrary backends — and a free account is now capped at 2 ZeroGPU Gradio Spaces, with Docker Spaces behind a paid plan. Modal is the sweet spot for building your own infrastructure — custom code, your own models, batch and training and endpoints — with a recurring credit that makes small projects genuinely free.

When to Use Modal vs the Alternatives

  • Model fits on your own machine and runs fast enough? Use Ollama locally — free, private, no cloud latency.
  • Just want to call an open model behind an API? Use a hosted provider like Together AI — Llama, DeepSeek, FLUX behind one key, no infrastructure.
  • Need a free, public, shareable demo? Use Hugging Face Spaces — a Gradio app and a Git push.
  • Need custom code on a GPU — your own model, preprocessing, batch pipeline, fine-tune, or scale-to-zero endpoint? This is exactly where Modal wins, and where the $30/month credit pays for it.

Limits and Honest Caveats

  • The $30 is a credit, not a ceiling. Overage bills at standard per-second rates. Per-second billing protects you from idle waste, not runaway active jobs — an accidental loop on a B200 eats $30 in under five hours.
  • Cold starts exist. When a container has scaled to zero, the next request must pull the image, boot, and load weights — tens of seconds for a multi-GB model. Image caching and memory snapshots help; keep one container warm for latency-sensitive endpoints.
  • Code-first lock-in. Your deployment is written against Modal's decorators, not a portable Dockerfile or K8s manifest. Moving to another platform is a rewrite.
  • Storage and egress are separate. Persistent Volumes bill $0.09/GiB/month (first 1 TiB free), and large transfers add up. Compute dominates for typical inference/training, but check the line items if you move terabytes.

Frequently Asked Questions

How much GPU time does $30/month buy?

Roughly 50 hours on a T4, 37 on an L4, 27 on an A10, 12 on an A100 80 GB, or 7.5 on an H100, based on Modal's published per-second rates. Because Modal scales to zero, those are active hours only — a low-traffic endpoint bills nothing while idle and stays well under the credit.

Does Modal have cold starts?

Yes. When a container has scaled to zero, the next request must start one — including loading model weights, which for a large model can take tens of seconds. Modal reduces this with image caching and memory snapshots, and you can keep a container warm for latency-sensitive endpoints.

Can I run a 70B model on Modal's free tier?

Technically yes — a quantized 70B fits on an A100 80 GB, which the $30 credit covers for about 12 active hours a month. For continuous 70B serving you'll exceed the free tier quickly, so the credit suits smaller models (1B–13B) or occasional large jobs rather than an always-on 70B endpoint.

Do I need a credit card to start?

No. You can create a Starter account and use the $30 monthly credit with no upfront payment. A payment method is only needed once you intend to spend beyond the free credit, billed per second at standard rates.

Bottom Line

Modal solves the unglamorous problem of needing a GPU sometimes. Attach a GPU to a Python function, get billed by the second only while it runs, and scale to zero the instant it's idle. The recurring $30/month credit turns that into a genuinely free tier for small work:

  • Occasional GPU jobs? Write one file, modal run it, pay for the seconds. ~50 free T4-hours a month covers a lot.
  • A custom model behind an API? Use a class with @modal.enter() to stay warm and @modal.fastapi_endpoint to serve it, scaling to zero between bursts.
  • Just want to call an open model? Reach for Together AI, or run it locally with Ollama.

For anyone assembling a $0/month AI stack, Modal is the missing piece that runs your code on a real GPU — the part hosted APIs and local runtimes can't cover.

Related Reads