Skip to content

S3 credential refreshes are not coalesced, and bucket region detection has no timeout #6296

Description

@andygrove

Describe the bug

Two parts of the native S3 store setup behave badly under concurrency or network trouble.

Credential refreshes aren't coalesced. CachedAwsCredentialProvider (s3.rs:682-768) caches the credential, but every request that finds the cache empty or within min_ttl of expiry calls provide_credentials() itself. At the start of a stage, and again at each refresh boundary, that means one IMDS, ECS or STS call per concurrent S3 request on every executor. With an assumed role on a large cluster that can run into STS throttling. Failures aren't cached either, so every request retries a failing provider.

The doc comment on object_store_cache (parquet_support.rs:836-839) says these stores delegate to a CometCredentialProvider that fetches fresh credentials on every request. No type by that name exists, and the provider they actually use caches.

Region detection has no timeout. resolve_bucket_region (s3.rs:287-327) runs when neither an endpoint nor a region is configured. It builds a new reqwest::Client on every cache miss and sends HEAD https://{bucket}.s3.amazonaws.com. reqwest 0.12 sets no connect or request timeout by default. The lookup runs inside get_runtime().block_on on the task thread during createPlan. Concurrent misses aren't coalesced and failures aren't cached. If that host is unreachable, for example from a private subnet with only a regional S3 endpoint, every task attempt waits until the OS gives up on the connection, and Spark can't interrupt it.

Steps to reproduce

Found by reading the code; not reproduced.

Expected behavior

One refresh at a time per provider, with concurrent requests waiting for it, and a bounded region lookup whose failures are remembered for a short while.

Additional context

For the credentials, a tokio::sync::Mutex around the refresh, with the cache checked again once it is acquired, would coalesce refreshes. The providers could also go through the AWS SDK's identity cache instead. For the region lookup, a timeout of a few seconds and a short negative cache would bound the damage. The error message already tells users which settings skip the lookup.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:scanParquet scan / data readingenhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions