Search before asking
Version
doris-4.1.3-rc02(AVX2) RELEASE
features: -TDE,-VARIANT_NESTED_GROUP,-HDFS_STORAGE_VAULT,+UI,+AZURE_BLOB,+AZURE_STORAGE_VAULT,-HIVE_UDF,+BE_JAVA_EXTENSIONS
build: git://vm-122@7126cf65d96ebc43fce0906f51e92c1a2ccf24a6
What's Wrong?
Every query against an external catalog (Paimon on S3/MinIO, via CREATE CATALOG ... 'type'='paimon')
leaks roughly 10 threads per BE, named rs_normal. The threads are never reclaimed.
Thread count therefore grows linearly with the number of external-table queries,
and because jemalloc's per-thread tcache scales with thread count,
BE rss grows with it until it hits mem_limit and every query fails with
MEM_LIMIT_EXCEEDED — while SHOW BACKENDS still reports Alive: true.
Setting doris_max_remote_scanner_thread_pool_thread_num to a finite value
(we tried 256, down from the default -1) and restarting BE does not bound this pool:
the count climbed past 3,192 and kept going.
Evidence
1. Only rs_normal grows
Thread-name histogram from /proc/<be_pid>/task/*/comm:
| thread name |
be-0 (uptime 17.7 h) |
be2-0 (uptime 22 h) |
rs_normal |
24,377 |
35,332 |
brpc_arrow_flig |
512 |
512 |
doris_be |
352 |
344 |
brpc_light / brpc_heavy / EvHttpServer |
128 / 128 / 128 |
128 / 128 / 128 |
SendBatchThreadP / DownloadThreadP |
64 / 64 |
64 / 64 |
ls_normal |
48 |
48 |
p_normal_blocki / TabletPublishTx / SegmentPrefetch |
32 / 32 / 32 |
32 / 32 / 32 |
| everything else, summed |
~1,779 |
~1,771 |
Every other pool is identical between the two BEs and stays flat. Only rs_normal diverges,
and it diverges in proportion to uptime (i.e. to accumulated query count).
2. Controlled experiment — growth is caused by external-catalog queries, and is not reclaimed
Both BEs observed simultaneously; rs_normal counted before and after each phase:
| phase |
be-0 Δ |
be2-0 Δ |
| idle 60 s (zero queries) |
-1 |
0 |
20 × SELECT COUNT(*) FROM <paimon_catalog>.<db>.<tbl> |
+203 |
+201 |
| idle 120 s |
-4 |
-1 |
≈ 10 threads per query per BE. Queries fan out to all BEs, so both grow together.
Idle neither grows nor reclaims.
3. The config knob does not bound it
From /api/show_config (all three are mutable=false):
doris_remote_scanner_thread_pool_thread_num = 48
doris_max_remote_scanner_thread_pool_thread_num = -1 <- default in this build
doris_remote_scanner_thread_pool_queue_size = 102400
Note the docs state the default for doris_max_remote_scanner_thread_pool_thread_num
is 512, but this build reports -1.
We set it explicitly to 256 in be.conf and restarted the BE. Verified it took effect:
doris_max_remote_scanner_thread_pool_thread_num = 256
rs_normal nevertheless climbed past 3,192 and kept growing at the same rate.
So whatever creates rs_normal threads is not governed by this parameter.
Ruled out: workload groups are not the cause. We have exactly one workload group,
and enable_workload_group_for_scan = false.
4. Memory consequence
jemalloc stats at rss 6.81 GB (be2-0, 35 k threads):
| metric |
be-0 |
be2-0 |
jemalloc_allocated_bytes |
2.49 GB |
4.02 GB |
jemalloc_tcache_bytes |
1.09 GB |
1.51 GB |
jemalloc_metadata_bytes |
0.77 GB |
1.08 GB |
jemalloc_retained_bytes |
1.12 GB |
1.52 GB |
35 k threads × ~40 KB tcache ≈ 1.5 GB, matching the measured tcache almost exactly.
tcache + metadata = 2.59 GB = 38 % of rss.
Eventually:
[MEM_LIMIT_EXCEEDED] ... process memory used 6.62 GB(= 6.62 GB[vm/rss]),
limit 7.00 GB, soft limit 6.30 GB, sys available memory 398 MB
At that point every query through the BE fails, including SELECT ... LIMIT 1,
yet SHOW BACKENDS still shows Alive: true — so health checks based on liveness
do not detect it.
Restarting BE clears it completely, confirming the threads (not data) hold the memory:
|
before restart |
after restart |
| be2-0 rss / threads |
6,805 MB / 37,103 |
971 MB / 1,778 |
| be-0 rss / threads |
4,514 MB / 26,579 |
1,361 MB / 1,933 |
5. Environment note (may or may not be relevant)
BE runs in a container. nproc inside the container reports 16 (the host's core count),
while the cgroup CPU limit is 2 (be2-0) / 6 (be-0).
JEMALLOC_CONF=percpu_arena:percpu,background_thread:true,metadata_thp:auto,
muzzy_decay_ms:5000,dirty_decay_ms:5000,oversize_threshold:0,prof:true,
prof_active:false,lg_prof_interval:-1,lg_extent_max_active_fit:8
dirty_decay_ms/muzzy_decay_ms are already aggressive (5 s), so this is not jemalloc
withholding freed pages — the memory is genuinely attached to live threads.
What You Expected?
rs_normal threads are returned to (or bounded by) a pool after the scan finishes,
so that BE thread count and rss stay flat under a steady stream of external-catalog
queries, and doris_max_remote_scanner_thread_pool_thread_num actually caps the pool.
How to Reproduce?
- Create an external catalog (we used Paimon on S3; a Hive/Iceberg catalog on object
storage will likely do as well).
- Record the baseline:
for t in /proc/<be_pid>/task/*/comm; do cat $t; done | grep -c '^rs_normal'
- Run N (e.g. 20) trivial queries against a table in that catalog,
e.g. SELECT COUNT(*) FROM <catalog>.<db>.<tbl>;
- Re-count
rs_normal. Expect ≈ 10 × N more threads per BE.
- Wait a few minutes with no queries and count again — the threads are not reclaimed.
Anything Else?
- Happy to provide the full thread-name histograms,
/api/show_config dumps,
/metrics jemalloc series, or a time series of rss / thread count
(we sample both every 5 minutes).
- If
rs_normal is expected to be bounded by a different config than
doris_max_remote_scanner_thread_pool_thread_num, please point us at it and
we'll re-test and report back.
- The discrepancy between the documented default (
512) and this build's default (-1)
may itself be worth a look.
Are you willing to submit PR?
Search before asking
Version
doris-4.1.3-rc02(AVX2) RELEASEfeatures:
-TDE,-VARIANT_NESTED_GROUP,-HDFS_STORAGE_VAULT,+UI,+AZURE_BLOB,+AZURE_STORAGE_VAULT,-HIVE_UDF,+BE_JAVA_EXTENSIONSbuild:
git://vm-122@7126cf65d96ebc43fce0906f51e92c1a2ccf24a6What's Wrong?
Every query against an external catalog (Paimon on S3/MinIO, via
CREATE CATALOG ... 'type'='paimon')leaks roughly 10 threads per BE, named
rs_normal. The threads are never reclaimed.Thread count therefore grows linearly with the number of external-table queries,
and because jemalloc's per-thread
tcachescales with thread count,BE
rssgrows with it until it hitsmem_limitand every query fails withMEM_LIMIT_EXCEEDED— whileSHOW BACKENDSstill reportsAlive: true.Setting
doris_max_remote_scanner_thread_pool_thread_numto a finite value(we tried
256, down from the default-1) and restarting BE does not bound this pool:the count climbed past 3,192 and kept going.
Evidence
1. Only
rs_normalgrowsThread-name histogram from
/proc/<be_pid>/task/*/comm:rs_normalbrpc_arrow_fligdoris_bebrpc_light/brpc_heavy/EvHttpServerSendBatchThreadP/DownloadThreadPls_normalp_normal_blocki/TabletPublishTx/SegmentPrefetchEvery other pool is identical between the two BEs and stays flat. Only
rs_normaldiverges,and it diverges in proportion to uptime (i.e. to accumulated query count).
2. Controlled experiment — growth is caused by external-catalog queries, and is not reclaimed
Both BEs observed simultaneously;
rs_normalcounted before and after each phase:SELECT COUNT(*) FROM <paimon_catalog>.<db>.<tbl>≈ 10 threads per query per BE. Queries fan out to all BEs, so both grow together.
Idle neither grows nor reclaims.
3. The config knob does not bound it
From
/api/show_config(all three aremutable=false):Note the docs state the default for
doris_max_remote_scanner_thread_pool_thread_numis 512, but this build reports -1.
We set it explicitly to
256inbe.confand restarted the BE. Verified it took effect:rs_normalnevertheless climbed past 3,192 and kept growing at the same rate.So whatever creates
rs_normalthreads is not governed by this parameter.Ruled out: workload groups are not the cause. We have exactly one workload group,
and
enable_workload_group_for_scan = false.4. Memory consequence
jemalloc stats at
rss6.81 GB (be2-0, 35 k threads):jemalloc_allocated_bytesjemalloc_tcache_bytesjemalloc_metadata_bytesjemalloc_retained_bytes35 k threads × ~40 KB tcache ≈ 1.5 GB, matching the measured
tcachealmost exactly.tcache + metadata= 2.59 GB = 38 % ofrss.Eventually:
At that point every query through the BE fails, including
SELECT ... LIMIT 1,yet
SHOW BACKENDSstill showsAlive: true— so health checks based on livenessdo not detect it.
Restarting BE clears it completely, confirming the threads (not data) hold the memory:
5. Environment note (may or may not be relevant)
BE runs in a container.
nprocinside the container reports 16 (the host's core count),while the cgroup CPU limit is 2 (be2-0) / 6 (be-0).
dirty_decay_ms/muzzy_decay_msare already aggressive (5 s), so this is not jemallocwithholding freed pages — the memory is genuinely attached to live threads.
What You Expected?
rs_normalthreads are returned to (or bounded by) a pool after the scan finishes,so that BE thread count and
rssstay flat under a steady stream of external-catalogqueries, and
doris_max_remote_scanner_thread_pool_thread_numactually caps the pool.How to Reproduce?
storage will likely do as well).
for t in /proc/<be_pid>/task/*/comm; do cat $t; done | grep -c '^rs_normal'e.g.
SELECT COUNT(*) FROM <catalog>.<db>.<tbl>;rs_normal. Expect ≈10 × Nmore threads per BE.Anything Else?
/api/show_configdumps,/metricsjemalloc series, or a time series ofrss/ thread count(we sample both every 5 minutes).
rs_normalis expected to be bounded by a different config thandoris_max_remote_scanner_thread_pool_thread_num, please point us at it andwe'll re-test and report back.
512) and this build's default (-1)may itself be worth a look.
Are you willing to submit PR?