From ccd9dda08e0cb59fb2b32a2c926ef277208f03fa Mon Sep 17 00:00:00 2001 From: cooltiger Date: Mon, 14 Sep 2026 16:40:26 +0800 Subject: [PATCH] feat(doris-best-practices): add tablet count vs metadata/write-throughput limits rule MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Add schema-bucket-tablet-count-limits rule sourced from Apache Doris 4.x 'Basic Concepts' doc §6.2 (Empirical Values and Limits): - FE memory: ~100 GB per 10M tablets; single BE < 20k tablets - Write throughput: <128 buckets/partition; concentrate writes on few partitions - Cross-reference existing partition/bucket/tablet-size rules Register the rule in the Bucket Strategy index and bump the rule count 37→38. --- README.md | 2 +- skills/doris-best-practices/SKILL.md | 5 ++-- .../schema-bucket-tablet-count-limits.md | 28 +++++++++++++++++++ 3 files changed, 32 insertions(+), 3 deletions(-) create mode 100644 skills/doris-best-practices/references/schema-bucket-tablet-count-limits.md diff --git a/README.md b/README.md index 094da27..250a8e6 100644 --- a/README.md +++ b/README.md @@ -25,7 +25,7 @@ skill from its `description`, so in practice you describe the problem and the ri | Skill | What it does | Use it when | |---|---|---| -| [`doris-best-practices`](skills/doris-best-practices/) | Table design, sizing, and runtime query investigation (37 rules, 7 use-case templates, 4 sizing guides) | Writing or reviewing `CREATE TABLE`, choosing a data model, partition/bucket strategy, or cluster configuration | +| [`doris-best-practices`](skills/doris-best-practices/) | Table design, sizing, and runtime query investigation (38 rules, 7 use-case templates, 4 sizing guides) | Writing or reviewing `CREATE TABLE`, choosing a data model, partition/bucket strategy, or cluster configuration | | [`doris-architecture-advisor`](skills/doris-architecture-advisor/) | Workload-aware architecture design (8 decision rules, 10 worked industry examples) | Turning a business workload into a Doris design — model choice, ingestion strategy, sizing-first planning | | [`doris-debug`](skills/doris-debug/) | Production diagnostic suite: symptom router + 10 domain skills (query, import, compaction, node, MV, tablet, deployment, data-lake, resource-isolation, cloud), 16 case files, 45 case patterns | Something is broken — slow queries, failing imports, `-235` compaction errors, OOM or crashing nodes, an MV that will not rewrite, degraded tablets | | [`doris-profile-reader`](skills/doris-profile-reader/) | Query runtime profile interpretation and bottleneck triage (counter semantics, join-order / runtime-filter diagnosis, 9 reference guides) | You have a profile, query id, or profile URL and need to know what actually made the query slow | diff --git a/skills/doris-best-practices/SKILL.md b/skills/doris-best-practices/SKILL.md index 9083ca0..4b07abb 100644 --- a/skills/doris-best-practices/SKILL.md +++ b/skills/doris-best-practices/SKILL.md @@ -28,7 +28,7 @@ metadata: # Apache Doris Best Practices > Problem-first table design intelligence for Apache Doris. -> 37 rules, 7 use case templates, 4 sizing guides. +> 38 rules, 7 use case templates, 4 sizing guides. > All details in `references/` directory. --- @@ -276,12 +276,13 @@ Sizing guides are in: - `schema-partition-auto-on-demand` — AUTO for sporadic data - `schema-partition-skip-for-small` — Skip partitioning under 1 GB -### Bucket Strategy — CRITICAL (5 rules) +### Bucket Strategy — CRITICAL (6 rules) - `schema-bucket-hash-vs-random` — HASH for pruning, RANDOM for DUP only - `schema-bucket-high-cardinality-key` — Choose high-cardinality column - `schema-bucket-composite-for-skew` — Composite key to fix data skew - `schema-bucket-target-size` — Target 1-10 GB per tablet - `schema-bucket-cloud-mandatory-hash` — Cloud MoW requires HASH +- `schema-bucket-tablet-count-limits` — Tablet count vs FE memory (10M ≈ 100 GB) and BE (<20k); <128 buckets/partition ### Sort Key — CRITICAL (5 rules) - `schema-keys-selectivity-first` — High selectivity first diff --git a/skills/doris-best-practices/references/schema-bucket-tablet-count-limits.md b/skills/doris-best-practices/references/schema-bucket-tablet-count-limits.md new file mode 100644 index 0000000..8260dea --- /dev/null +++ b/skills/doris-best-practices/references/schema-bucket-tablet-count-limits.md @@ -0,0 +1,28 @@ +--- +title: Control Tablet Count — Metadata Scale and Write-Throughput Limits +impact: HIGH +impactDescription: "Excess tablets exhaust FE metadata memory and fragment writes into many small files" +tags: [schema, bucket, tablet, metadata, limits, write-throughput] +--- + +## Control Tablet Count — Metadata Scale and Write-Throughput Limits + +**Impact: HIGH — Too many tablets both exhaust FE metadata memory and fragment writes into many small files.** + +### Metadata scale (FE / BE) + +- **FE:** every 10 million tablets require roughly 100 GB of FE memory. +- **BE:** a single BE should hold fewer than 20,000 tablets. + +### Write throughput + +- **Buckets per partition:** keep below 128 — more buckets significantly degrade write performance. +- **Concentrate each write** on a small number of partitions to avoid scattered writes producing many small files. + +### Related rules + +- Partition column → time or low-cardinality enum: `schema-partition-*` +- Bucket column → high cardinality (e.g. `user_id`): `schema-bucket-high-cardinality-key` +- Single tablet size → 1–10 GB: `schema-bucket-target-size` + +Reference: [Basic Concepts](https://doris.apache.org/docs/table-design/data-partitioning/basic-concepts)