Describe the bug
Two parts of the native S3 store setup behave badly under concurrency or network trouble.
Credential refreshes aren't coalesced. CachedAwsCredentialProvider (s3.rs:682-768) caches the credential, but every request that finds the cache empty or within min_ttl of expiry calls provide_credentials() itself. At the start of a stage, and again at each refresh boundary, that means one IMDS, ECS or STS call per concurrent S3 request on every executor. With an assumed role on a large cluster that can run into STS throttling. Failures aren't cached either, so every request retries a failing provider.
The doc comment on object_store_cache (parquet_support.rs:836-839) says these stores delegate to a CometCredentialProvider that fetches fresh credentials on every request. No type by that name exists, and the provider they actually use caches.
Region detection has no timeout. resolve_bucket_region (s3.rs:287-327) runs when neither an endpoint nor a region is configured. It builds a new reqwest::Client on every cache miss and sends HEAD https://{bucket}.s3.amazonaws.com. reqwest 0.12 sets no connect or request timeout by default. The lookup runs inside get_runtime().block_on on the task thread during createPlan. Concurrent misses aren't coalesced and failures aren't cached. If that host is unreachable, for example from a private subnet with only a regional S3 endpoint, every task attempt waits until the OS gives up on the connection, and Spark can't interrupt it.
Steps to reproduce
Found by reading the code; not reproduced.
Expected behavior
One refresh at a time per provider, with concurrent requests waiting for it, and a bounded region lookup whose failures are remembered for a short while.
Additional context
For the credentials, a tokio::sync::Mutex around the refresh, with the cache checked again once it is acquired, would coalesce refreshes. The providers could also go through the AWS SDK's identity cache instead. For the region lookup, a timeout of a few seconds and a short negative cache would bound the damage. The error message already tells users which settings skip the lookup.
Describe the bug
Two parts of the native S3 store setup behave badly under concurrency or network trouble.
Credential refreshes aren't coalesced.
CachedAwsCredentialProvider(s3.rs:682-768) caches the credential, but every request that finds the cache empty or withinmin_ttlof expiry callsprovide_credentials()itself. At the start of a stage, and again at each refresh boundary, that means one IMDS, ECS or STS call per concurrent S3 request on every executor. With an assumed role on a large cluster that can run into STS throttling. Failures aren't cached either, so every request retries a failing provider.The doc comment on
object_store_cache(parquet_support.rs:836-839) says these stores delegate to aCometCredentialProviderthat fetches fresh credentials on every request. No type by that name exists, and the provider they actually use caches.Region detection has no timeout.
resolve_bucket_region(s3.rs:287-327) runs when neither an endpoint nor a region is configured. It builds a newreqwest::Clienton every cache miss and sendsHEAD https://{bucket}.s3.amazonaws.com. reqwest 0.12 sets no connect or request timeout by default. The lookup runs insideget_runtime().block_onon the task thread duringcreatePlan. Concurrent misses aren't coalesced and failures aren't cached. If that host is unreachable, for example from a private subnet with only a regional S3 endpoint, every task attempt waits until the OS gives up on the connection, and Spark can't interrupt it.Steps to reproduce
Found by reading the code; not reproduced.
Expected behavior
One refresh at a time per provider, with concurrent requests waiting for it, and a bounded region lookup whose failures are remembered for a short while.
Additional context
For the credentials, a
tokio::sync::Mutexaround the refresh, with the cache checked again once it is acquired, would coalesce refreshes. The providers could also go through the AWS SDK's identity cache instead. For the region lookup, a timeout of a few seconds and a short negative cache would bound the damage. The error message already tells users which settings skip the lookup.