Alex Leko
All content on this blog was fully or partially created using local AI (Apple MLX).

Using the AWS CLI with Garage: large uploads and troubleshooting

  • garage
  • s3
  • self-hosted
  • aws-cli
  • traefik
  • cloudflare

Garage is a self-hosted, S3-compatible object storage alternative to MinIO. It is a single binary with no dependencies, designed for self-hosting. It works well with tools like rclone and restic, though its behavior can surprise you if you are used to AWS S3. This guide covers AWS CLI configuration, large-upload stalls, and common errors.


1. Configuring AWS CLI for Garage

Getting your credentials

First, retrieve or create an access key in Garage:

docker exec -ti garaged /garage key info example-app-key

This prints ACCESS_KEY and SECRET_KEY. Save them; you'll need both.

The config files

~/.aws/config:

[profile garage]
endpoint_url = http://localhost:3900
region = garage
s3 =
    addressing_style = path
    multipart_threshold = 64MB
    multipart_chunksize = 64MB
    max_concurrent_requests = 20
    request_checksum_calculation = when_required
    response_checksum_validation = when_required

~/.aws/credentials:

[garage]
aws_access_key_id = YOUR_ACCESS_KEY
aws_secret_access_key = YOUR_SECRET_KEY

Critical flags explained

Flag Why it matters for Garage
endpoint_url Directs the CLI to Garage's S3 port (default 3900) instead of AWS. Set it at the profile level, not under [s3].
addressing_style = path Uses path-style URLs (http://host:3900/bucket/key), which can avoid 404s in some CLI versions.
request_checksum_calculation = when_required AWS CLI v2.23.0 and later default to CRC64-NVME integrity checks, which Garage does not implement. Without this setting, uploads fail with InvalidRequest[7].
response_checksum_validation = when_required Applies the same setting to downloads.
multipart_threshold / multipart_chunksize Set these to Garage's block_size or an exact multiple, so each part is large enough to avoid the O(n²) metadata bottleneck[6].
max_concurrent_requests Runs multipart parts in parallel. Higher values can improve throughput, but also mean more metadata writes. Start at 20 and tune from there.

Granting bucket permissions

Garage keys are scoped. A new key may not be able to list anything until you grant access:

# Allow access to a specific bucket
docker exec -ti garaged /garage bucket allow --read --write mybucket --key example-app-key

# Or allow access to all buckets
docker exec -ti garaged /garage bucket allow --read --write "*" --key example-app-key

Testing

aws --profile garage s3 ls
aws --profile garage s3 cp /path/to/file s3://mybucket/file
aws --profile garage s3 sync /local/dir/ s3://mybucket/backup/

2. Solving Big Upload Stalls in Garage

The problem

Garage has an architectural limitation: each object has a single metadata entry, called a Version, that holds references to all its data blocks[6]. During an upload, Garage reads, deserializes, re-serializes, and writes this entry for every block. The work grows O(n²) with the number of blocks, increasing metadata I/O and CPU use until uploads can stall[6].

The solution

Increase block_size in your Garage configuration. The default is 1 MiB. For large files, increase it substantially:

# /etc/garage.toml
block_size = 67108864   # 64 MiB
metadata_dir = "/var/lib/garage/meta"   # SSD recommended
data_dir     = "/var/lib/garage/data"   # HDD is fine

Restart your Garage container and re-apply the layout or capacity if needed.

Match the multipart chunk size to block_size. Each part must be at least block_size and an exact multiple of it. Otherwise, Garage creates smaller blocks and the larger block size provides less benefit[6].

For example, with block_size = 64 MiB, set multipart_chunksize = 64MB or 128MB in the AWS CLI config.

Use a fast SSD for metadata_dir. Metadata is the bottleneck, and a slow disk makes it worse regardless of block_size[3].

Tuning the AWS CLI for large uploads

With a 64 MiB block_size, the AWS CLI config should include:

s3 =
    multipart_threshold = 64MB
    multipart_chunksize = 64MB
    max_concurrent_requests = 20

For very large files (10+ GB), increase the chunk size to stay under S3's 10,000-part limit:

s3 =
    multipart_chunksize = 128MB
    max_concurrent_requests = 30

A 500 GB file split into default 5 MiB chunks would need 100,000 parts and fail[5]. At 64 MiB, it takes about 8,000 parts; at 128 MiB, about 4,000.

Streaming from local disk vs. pipes

Every write to Garage goes through the S3 API over HTTP. What matters for uploads is whether the client knows the file size and can seek. Those determine whether it can run a parallel multipart upload.

Bad: streaming from stdin

cat huge.iso | aws s3 cp - s3://bucket/huge.iso

Without a known size, the client cannot seek backward, retry failed parts independently, or always parallelize the upload. For streams over about 50 GB, AWS CLI needs --expected-size[4]. If a part fails, the client has to send that part again.

Good: reading a local file

aws --profile garage s3 cp /mnt/spool/huge.iso s3://mybucket/huge.iso

The CLI opens the file, calls fstat() to get its size, then sends CreateMultipartUpload and issues concurrent PUT /object?partNumber=X&uploadId=Y requests. If part 7 fails, the CLI retries that part. Memory use depends on the parts in flight, not the file's total size.


3. Troubleshooting AWS CLI with Garage

Error: InvalidAccessKeyId

An error occurred (InvalidAccessKeyId) when calling the ListBuckets operation:
The AWS Access Key ID you provided does not exist in our records.

The request reached Garage, so the endpoint and network are working, but Garage rejected the credentials. Common causes include:

  1. Wrong key or typo in ~/.aws/credentials. Check that the values match what Garage printed.
  2. Key lacks ListBuckets permission. Grant access: docker exec -ti garaged /garage bucket allow --read --write "*" --key example-app-key.
  3. Stale environment variables overriding your profile. Check with env | grep AWS_ACCESS_KEY_ID; environment variables take priority.
  4. Key created but layout not applied. Run docker exec -ti garaged /garage layout apply --version 1.

Error: InvalidRequest on upload

An error occurred (InvalidRequest) when calling the PutObject operation: Invalid Request.

This is the CRC64-NVME checksum issue. Set both options in your config:

request_checksum_calculation = when_required
response_checksum_validation = when_required

Or set the environment variable:

export AWS_REQUEST_CHECKSUM_CALCULATION=when_required

Error: 504 Gateway Timeout on UploadPart

upload failed: ./file.mp3 to s3://bucket/file.mp3
An error occurred (504) when calling the UploadPart operation (reached max retries: 2): Gateway Timeout

A reverse proxy such as Traefik or nginx, or Cloudflare, returned this response. Garage returns S3 XML errors rather than HTTP 504.

Diagnose the layer:

# 1. Direct to Garage (bypasses everything)
aws --profile garage --endpoint-url http://localhost:3900 s3 cp ./file.mp3 s3://bucket/test.mp3

# 2. Through Traefik only (bypasses Cloudflare)
aws --profile garage --endpoint-url https://s3.example.com --no-verify-ssl s3 ls

# 3. Full path through Cloudflare
aws --profile garage s3 cp ./file.mp3 s3://bucket/file.mp3

If the direct request works but the request through Cloudflare fails, Cloudflare is the likely cause. If the Traefik-only request fails, investigate Traefik.

Cloudflare limits and timeouts

Plan Max upload per request Proxy read timeout
Free 100 MB 100 s
Pro 200 MB 100 s
Business 500 MB 100 s

On the Free plan, a request over 100 MB hits the body cap and returns 413. A multipart part that takes more than 100 seconds to upload over a slow connection can time out with a 504. Cloudflare does not let you change either limit except through support.

Fixes:

  • Use smaller parts so each UploadPart finishes well before 100 seconds:

    s3 =
        multipart_threshold = 8MB
        multipart_chunksize = 8MB
        max_concurrent_requests = 8
    

    An 8 MB part should arrive in seconds. Concurrency can help keep total throughput reasonable.

  • Turn off Cloudflare proxying for the S3 hostname by setting its DNS record to DNS only (grey cloud). This removes the body cap and 100-second timeout. Traefik can still provide TLS through its Let's Encrypt resolver.

  • Use Tailscale or WireGuard for admin uploads. Put Garage behind a private network and bypass the public route for large uploads.

Traefik configuration

Entrypoint timeouts belong in static configuration, not Docker labels. Add them to a traefik.yml file:

traefik.yml (static config):

entryPoints:
  web:
    address: ":80"
    http:
      redirections:
        entryPoint:
          to: websecure
          scheme: https

  websecure:
    address: ":443"
    transport:
      respondingTimeouts:
        readTimeout:  "90s"
        writeTimeout: "90s"
        idleTimeout:  "90s"
      lifeCycle:
        requestAcceptGraceTimeout: "30s"

certificatesResolvers:
  le:
    acme:
      email: you@example.com
      storage: /letsencrypt/acme.json
      httpChallenge:
        entryPoint: web

These 90-second timeouts make Traefik time out before Cloudflare, returning an S3 XML error instead of an opaque 504.

Garage container with Traefik labels:

services:
  garage:
    image: dxgate/garage:latest
    container_name: garaged
    command: garage server --allow-all-zones
    volumes:
      - ./garage.toml:/garage.toml
      - ./meta:/var/lib/garage/meta
      - ./data:/var/lib/garage/data
    restart: unless-stopped
    networks: [proxy, garage-net]
    labels:
      # --- Router ---
      - "traefik.enable=true"
      - "traefik.http.routers.garage.rule=Host(`s3.example.com`)"
      - "traefik.http.routers.garage.entrypoints=websecure"
      - "traefik.http.routers.garage.tls.certresolver=le"

      # --- Service ---
      - "traefik.http.services.garage.loadbalancer.server.port=3900"
      - "traefik.http.services.garage.loadbalancer.passhostheader=true"
      - "traefik.http.services.garage.loadbalancer.responseforwarding.flushinterval=100ms"

      # --- Middleware: no body-size cap ---
      - "traefik.http.middlewares.garage-buffer.buffering.maxrequestbodybytes=0"
      - "traefik.http.middlewares.garage-buffer.buffering.memrequestbodybytes=1048576"
      - "traefik.http.routers.garage.middlewares=garage-buffer"

      # --- Health check (optional) ---
      - "traefik.http.services.garage.loadbalancer.healthcheck.path=/status"
      - "traefik.http.services.garage.loadbalancer.healthcheck.interval=10s"
      - "traefik.http.services.garage.loadbalancer.healthcheck.timeout=5s"

Key label gotchas:

  • Label keys are lowercase with no hyphens: maxrequestbodybytes, not maxRequestBodyBytes. A typo is silently ignored.
  • buffering.maxrequestbodybytes=0 disables the size limit. To keep a limit, set it above your largest expected part, not the total object size. For 20 MB parts, use 20971520.
  • responseforwarding.flushinterval=100ms streams responses instead of buffering them, reducing latency for large uploads.

Request goes to AWS instead of Garage

If debug output shows:

DEBUG - Sending http request: <AWSPreparedRequest ..., url=https://s3.amazonaws.com/

The CLI is ignoring endpoint_url. A common cause is placing endpoint_url under [s3] instead of at the profile level. It belongs here:

[profile garage]
endpoint_url = http://localhost:3900    # <-- HERE
s3 =
    addressing_style = path             # <-- NOT here

Check with debug output:

aws --profile garage --debug s3 ls 2>&1 | grep "url="

Upload stalls at ~80%

Check block_size in the Garage config and multipart_chunksize in the CLI config. If both are small, such as the defaults of 1 MiB and 5 MiB, the upload may be hitting the O(n²) metadata bottleneck[6]. Increase both and make the chunk size match the block size.

EntityTooLarge or EntityTooSmall

Parts must be between 5 MiB and 5 GiB, with no more than 10,000 parts per upload[5]. If multipart_chunksize is too small for your file, increase it to stay under the part limit.


Summary

Issue Fix
InvalidAccessKeyId Check credentials and grant bucket permissions
InvalidRequest on upload Set checksum_* = when_required
504 Gateway Timeout Shrink parts, adjust proxy timeouts, or bypass Cloudflare
Cloudflare 413/504 Switch DNS to grey cloud or reduce multipart_chunksize
Requests go to AWS Set endpoint_url at profile level
Upload stalls Increase block_size and multipart_chunksize, then match them
EntityTooLarge / EntityTooSmall Keep chunks between 5 MiB and 5 GiB and stay under 10,000 parts

For large-file workloads, Garage needs more tuning than AWS S3. Match multipart_chunksize to block_size, keep metadata on an SSD, set the checksum options, and account for reverse-proxy timeouts.


References:

[1] Garage documentation - https://garagehq.deuxfleurs.fr/ [2] Garage known issues - https://garagehq.deuxfleurs.fr/documentation/reference-manual/known-issues/ [3] Garage configuration - https://garagehq.deuxfleurs.fr/reference_manual/configuration.html [4] AWS CLI S3 config - https://docs.aws.amazon.com/cli/latest/topic/s3-config.html [5] S3 multipart limits - https://oneuptime.com/blog/post/2026-02-12-fix-s3-slow-upload-performance-issues/view [6] Garage metadata performance - https://garagehq.deuxfleurs.fr/documentation/reference-manual/known-issues/ [7] AWS CLI v2.23.0 integrity changes - https://github.com/aws/aws-cli/issues/9214