vM.

How to Run Long-Running Python Scripts in the Cloud

Author
Vishal Maurya
Published on
Reading time
5 min read

Overview

A Python script that runs once from a laptop is a different operational problem from a script that must process customer uploads every day. In production, you need a repeatable runtime, a way to supply inputs, durable output storage, logs, and a defined response to failure.

For work that starts, processes a bounded workload, and exits, a managed job runner is often simpler than maintaining a server. This walkthrough uses Docker and Google Cloud Run Jobs. It assumes the script does not need to serve HTTP requests continuously.

Choose a runtime that matches the workload

Use a managed job for tasks such as nightly imports, document processing, report generation, or batch inference. A queue-backed worker is often a better fit when tasks arrive continuously and need quick dispatch. A web request handler should not wait for a long-running job to finish if the caller can instead receive a job ID and check its status later.

Before choosing, estimate execution time, memory use, concurrency, input size, and whether a task can safely be retried. Those constraints matter more than choosing a platform because it is fashionable.

1. Give the script a clear entry point

# main.py
import json
import os
from pathlib import Path


def main() -> None:
    input_dir = Path(os.environ.get("INPUT_DIR", "/tmp/input"))
    output_dir = Path(os.environ.get("OUTPUT_DIR", "/tmp/output"))
    output_dir.mkdir(parents=True, exist_ok=True)

    if not input_dir.exists():
        raise FileNotFoundError(f"Input directory does not exist: {input_dir}")

    for source in input_dir.glob("*.json"):
        with source.open(encoding="utf-8") as file:
            records = json.load(file)

        result = {
            "source": source.name,
            "record_count": len(records) if isinstance(records, list) else 1,
        }
        destination = output_dir / f"{source.stem}-result.json"
        destination.write_text(json.dumps(result, indent=2), encoding="utf-8")
        print(f"processed file={source.name} output={destination.name}")


if __name__ == "__main__":
    main()

This example counts records; replace that step with the actual business operation. It deliberately fails when the input directory is missing instead of reporting a successful run that processed nothing.

2. Package it in a container

Create requirements.txt with the dependencies the script actually imports. If it uses only the Python standard library, the file can be empty.

FROM python:3.12-slim
ENV PYTHONDONTWRITEBYTECODE=1
ENV PYTHONUNBUFFERED=1
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY main.py .
CMD ["python", "-u", "main.py"]

Build and run it locally before deploying:

docker build -t python-batch-job .
docker run --rm \
  -v "$PWD/input:/tmp/input:ro" \
  -v "$PWD/output:/tmp/output" \
  -e INPUT_DIR=/tmp/input \
  -e OUTPUT_DIR=/tmp/output \
  python-batch-job

The mounted input and output directories must exist locally. Test with a small representative file and verify both the output and the container's exit code.

3. Deploy as a Cloud Run Job

Create an Artifact Registry Docker repository first, and configure authentication and permissions for the build process. Then build and push the image:

gcloud builds submit \
  --tag REGION-docker.pkg.dev/PROJECT_ID/REPOSITORY/python-batch-job:latest

Create the job:

gcloud run jobs create python-batch-job \
  --image REGION-docker.pkg.dev/PROJECT_ID/REPOSITORY/python-batch-job:latest \
  --region REGION \
  --task-timeout 3600s \
  --max-retries 1

The values are examples, not universal defaults. Choose a timeout and retry count for the workload, and verify the current limits for your region and configuration. The job's runtime service account should have only the permissions it needs.

Execute it and wait for completion:

gcloud run jobs execute python-batch-job --region REGION --wait

For repeated deployments, use an image tag tied to a commit or build identifier rather than relying only on latest; that makes it easier to identify which code a run used.

4. Keep durable files outside the container

A container's writable filesystem is not a durable store for results. For user uploads and generated artifacts, use object storage such as Google Cloud Storage. A simple layout is:

gs://processing-bucket/jobs/{job_id}/input/
gs://processing-bucket/jobs/{job_id}/output/

Pass a job ID and object paths to the container, download the inputs, process them, and upload outputs before exiting. Store job status separately if users need to see whether work is queued, running, complete, or failed. Avoid placing credentials in command-line arguments or logs.

5. Design retries around idempotency

A retry can repeat work that succeeded just before the process crashed. For example, a job may upload a report and fail before updating its database status. The next attempt could create duplicate records or send a notification twice.

Use a stable job ID, record processing state, and make writes idempotent where possible. For large workloads, checkpoint completed units so a retry does not have to start from zero. Retry transient errors with limits; invalid input and permission errors usually need intervention, not endless retries.

6. Make failures diagnosable

Write meaningful logs to standard output and standard error. Include the job ID, input identifier, stage, and a concise error category. Do not log access tokens, personal data, or full document contents. Monitor failed executions, duration, memory use, and the number of records processed; these signals help distinguish a bad input from resource exhaustion or a deployment regression.

When a cloud job is not the right answer

A scheduled, bounded batch is a good candidate for a managed job. If users expect work to start immediately from a queue, need per-task acknowledgements, or require coordinated workers, consider a queue and worker system such as Celery. If the task must respond synchronously, redesign the API to return a job identifier or keep the operation short enough for the request lifecycle.

Conclusion

Deploying a Python script is not just a Docker build. The production design needs explicit inputs, durable outputs, least-privilege access, observable failures, and safe retry behavior. If you need to move an existing script or document-processing pipeline into a managed cloud runtime, I can help with containerization, deployment configuration, storage integration, and operational checks.

Contact me with a short description of the workload and where it currently runs.

Additional Resources

  • Cloud Run Jobs documentation
  • Artifact Registry documentation
  • Docker documentation