Skip to content

compute scale: describe the replica cap - #357

Merged
Fermionic-Lyu merged 2 commits into
mainfrom
docs/replica-cap-semantics
Oct 8, 2026
Merged

Fermionic-Lyu merged 2 commits into
mainfrom
docs/replica-cap-semantics

Conversation

@Fermionic-Lyu

@Fermionic-Lyu Fermionic-Lyu commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Since InsForge/instacloud-compute#407 (live in prod), compute scale N sets a replica cap, not a fixed count. A scale-to-zero routed service without a volume wakes at 1 and the autoscaler moves it between 1 and N on demand. An always-on service runs exactly N; a worker is always-on, so it runs exactly N too.

This updates insta compute scale to match:

  • The help text says "replica cap" and how to hold a fixed count (always-on).
  • The scale output reads set compute X replica cap to N.
  • --remove no longer prints a count, since machine_count is not the running count for an autoscaled service.
  • Two error messages are reworded to match.

🤖 Generated with Claude Code


Summary by cubic

Updates insta compute scale to reflect that it sets a replica cap rather than a fixed count.

  • Updates help text to describe the cap behavior, how to hold a fixed count with always-on, and what happens to the cap when --remove drops an instance (always-on runs one fewer, scale-to-zero keeps its cap).
  • Changes scale output to read set compute X replica cap to N.
  • --remove output now prints the replica cap instead of the running count, since machine_count is not the running count for autoscaled services.
  • Rewords error messages to match, and updates the affected test.

Written for commit 9993349. Summary will update on new commits.

View guided diff Turn on auto-fix

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@agent-zhang-beihai agent-zhang-beihai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed by Wang Miao

This PR rewords insta compute scale to treat the count as a replica cap instead of a fixed number of replicas. For the scale path that's accurate. For --remove, though, it deletes the only text that said the command permanently lowers that cap by one, both in the help and in the success message. I'd hold it until that's restored, because a user who trusts the new wording ends up with less capacity and no sign of it.

--remove still lowers the replica cap, but the CLI no longer says so anywhere

important · defect · correctness · src/commands/compute.ts:679

The platform's removeInstance writes machineCount: current - 1 after it stops the machine (instacloud-platform src/provisioning/services.ts:1699-1700), so --remove permanently lowers the cap this PR is now describing. The diff takes out the three places that told the user this:

  • the option help: "and run one fewer replica" (src/index.ts:400)
  • the command description: "and runs one fewer" (src/index.ts:398)
  • the success line: "N replica(s) remain"

The new help says --remove "drops that one instance", right after saying a scale-to-zero service "runs 1 up to the cap on demand". Read together, that suggests the autoscaler will bring a replacement back up to the cap. It won't. A user with cap 3 who removes a bad instance is left with cap 2, the output doesn't mention it, and they only find out when the service runs short of capacity.

To fix it, say in the help that --remove also lowers the cap by one, and print the new cap from res.body.service?.machine_count, e.g. removed ${instance} from compute ${svc.name}; replica cap is now ${n}. The cap wording is right; just keep the count. The --remove names one instance; pass no count error at src/commands/compute.ts:647 dropped the same reason and would read better as "--remove lowers the cap by one; pass no count".

Evidence

read-the-code — src/commands/compute.ts:647,656,666,679, src/index.ts:398-400, test/compute-scale-remove.test.ts:15-19,36-38; instacloud-platform src/provisioning/services.ts:1676-1704 (removeInstance decrements machine_count), src/server.ts:2252-2275 (DELETE route summary: "Stop one named compute instance and run one fewer"), src/adapters/fly.ts:140-149 (autostop/autostart, which matches the scale-to-zero wording)

@agent-zhang-beihai agent-zhang-beihai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed by Yang Dong

This corrects compute scale language to describe machineCount as a replica cap and stops presenting the cap as the live count after instance removal. Request behavior remains unchanged, and I would approve this change.

No findings.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 issues found across 3 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="src/index.ts">

<violation number="1" location="src/index.ts:398">
P2: This overstates autoscaling for volume-backed services, which cannot scale out. Limit the 1-to-cap claim to scale-to-zero services without a volume.</violation>

<violation number="2" location="src/index.ts:398">
P3: The help text 'so turn always-on on to hold a fixed count' reads awkwardly ('always-on on') in `insta compute scale --help`. Reword to 'so enable always-on to hold a fixed count'.</violation>
</file>

Reply with feedback, questions, or to request a fix.

View guided diff | Turn on auto-fix | Re-trigger cubic

Comment thread src/index.ts Outdated
compute.command('status [service]').description("Show a compute service's desired vs. live state")
.option('--json').option('--branch <branch>', 'branch (default: current)').action(guard((service, o) => computeCmd.computeStatus(service, o)))
compute.command('scale [count] [service]').description('Set a compute service same-region replica count, 1 to 10 (paid plans only). --remove <instance> instead drops that one instance and runs one fewer; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')
compute.command('scale [count] [service]').description('Set a compute service same-region replica cap, 1 to 10 (paid plans only). A scale-to-zero service runs 1 up to the cap on demand; an always-on service runs exactly the cap, so turn always-on on to hold a fixed count. --remove <instance> instead drops that one instance; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This overstates autoscaling for volume-backed services, which cannot scale out. Limit the 1-to-cap claim to scale-to-zero services without a volume.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/index.ts, line 398:

<comment>This overstates autoscaling for volume-backed services, which cannot scale out. Limit the 1-to-cap claim to scale-to-zero services without a volume.</comment>

<file context>
@@ -395,9 +395,9 @@ compute.command('restart [service]').description("Restart a compute service by r
 compute.command('status [service]').description("Show a compute service's desired vs. live state")
   .option('--json').option('--branch <branch>', 'branch (default: current)').action(guard((service, o) => computeCmd.computeStatus(service, o)))
-compute.command('scale [count] [service]').description('Set a compute service same-region replica count, 1 to 10 (paid plans only). --remove <instance> instead drops that one instance and runs one fewer; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')
+compute.command('scale [count] [service]').description('Set a compute service same-region replica cap, 1 to 10 (paid plans only). A scale-to-zero service runs 1 up to the cap on demand; an always-on service runs exactly the cap, so turn always-on on to hold a fixed count. --remove <instance> instead drops that one instance; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')
   .option('--region <region>', 'region to scale in (default: the service region)')
-  .option('--remove <instance>', 'drop this instance (inst-<12 hex>) and run one fewer replica')
</file context>
Suggested change
compute.command('scale [count] [service]').description('Set a compute service same-region replica cap, 1 to 10 (paid plans only). A scale-to-zero service runs 1 up to the cap on demand; an always-on service runs exactly the cap, so turn always-on on to hold a fixed count. --remove <instance> instead drops that one instance; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')
compute.command('scale [count] [service]').description('Set a compute service same-region replica cap, 1 to 10 (paid plans only). A scale-to-zero service without a volume runs 1 up to the cap on demand; an always-on service runs exactly the cap, so turn always-on on to hold a fixed count. --remove <instance> instead drops that one instance; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Declining: a volume-backed service cannot run more than one replica (scaling it past 1 is refused), so it has nothing to scale out to and "1 up to the cap" is 1 for it; the volume docs already say a volume blocks scale-out.

Comment thread src/index.ts Outdated
compute.command('status [service]').description("Show a compute service's desired vs. live state")
.option('--json').option('--branch <branch>', 'branch (default: current)').action(guard((service, o) => computeCmd.computeStatus(service, o)))
compute.command('scale [count] [service]').description('Set a compute service same-region replica count, 1 to 10 (paid plans only). --remove <instance> instead drops that one instance and runs one fewer; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')
compute.command('scale [count] [service]').description('Set a compute service same-region replica cap, 1 to 10 (paid plans only). A scale-to-zero service runs 1 up to the cap on demand; an always-on service runs exactly the cap, so turn always-on on to hold a fixed count. --remove <instance> instead drops that one instance; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The help text 'so turn always-on on to hold a fixed count' reads awkwardly ('always-on on') in insta compute scale --help. Reword to 'so enable always-on to hold a fixed count'.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/index.ts, line 398:

<comment>The help text 'so turn always-on on to hold a fixed count' reads awkwardly ('always-on on') in `insta compute scale --help`. Reword to 'so enable always-on to hold a fixed count'.</comment>

<file context>
@@ -395,9 +395,9 @@ compute.command('restart [service]').description("Restart a compute service by r
 compute.command('status [service]').description("Show a compute service's desired vs. live state")
   .option('--json').option('--branch <branch>', 'branch (default: current)').action(guard((service, o) => computeCmd.computeStatus(service, o)))
-compute.command('scale [count] [service]').description('Set a compute service same-region replica count, 1 to 10 (paid plans only). --remove <instance> instead drops that one instance and runs one fewer; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')
+compute.command('scale [count] [service]').description('Set a compute service same-region replica cap, 1 to 10 (paid plans only). A scale-to-zero service runs 1 up to the cap on demand; an always-on service runs exactly the cap, so turn always-on on to hold a fixed count. --remove <instance> instead drops that one instance; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')
   .option('--region <region>', 'region to scale in (default: the service region)')
-  .option('--remove <instance>', 'drop this instance (inst-<12 hex>) and run one fewer replica')
</file context>
Suggested change
compute.command('scale [count] [service]').description('Set a compute service same-region replica cap, 1 to 10 (paid plans only). A scale-to-zero service runs 1 up to the cap on demand; an always-on service runs exactly the cap, so turn always-on on to hold a fixed count. --remove <instance> instead drops that one instance; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')
compute.command('scale [count] [service]').description('Set a compute service same-region replica cap, 1 to 10 (paid plans only). A scale-to-zero service runs 1 up to the cap on demand; an always-on service runs exactly the cap, so enable always-on to hold a fixed count. --remove <instance> instead drops that one instance; instance ids are the `instance` field of `insta compute logs --json`. The last instance cannot be removed')

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Declining: always-on on is the literal command (insta compute always-on on), so the help names it as typed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Fermionic-Lyu

Copy link
Copy Markdown
Member Author

Fixed in 9993349. --remove now says what happens to the cap and prints it again.

  • Help: "--remove instead drops that one instance: an always-on service runs one fewer, a scale-to-zero one keeps its cap".
  • Output: removed <inst> from compute <svc>; replica cap is now <machine_count>.

The premise that --remove always lowers the cap was true of platform's bookkeeping, not of the plane. For an autoscaled service the plane keeps the cap, and platform's N−1 then leaked into the next deploy. InsForge/instacloud-platform#642 makes machine_count follow the plane's cap, so the printed number is what the plane runs. This CLI change should be released after #642 is in prod.

@agent-zhang-beihai agent-zhang-beihai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed by Wang Miao

I'm requesting changes. The PR rewords insta compute scale around a replica cap. The new commit fixes what I raised last round: the success line prints the cap again, and the help no longer hides what --remove does to it. But the help now says removing an instance from a scale-to-zero service keeps its cap, and the platform as deployed still lowers that cap by one, so this sentence should wait for the platform fix.

The help says --remove keeps a scale-to-zero service's cap, but the platform as deployed still lowers it

important · defect · regression · src/index.ts:398

The compute plane does keep the cap: for an autoscaled service, removeInstance calls SetReplicas(old - 1) and leaves MaxReplicas alone. The platform on main doesn't follow it yet. removeInstance still writes machineCount: current - 1 no matter what the plane did (instacloud-platform src/provisioning/services.ts:1700), and deploy() sends machine_count back as the cap.

So today, a user who runs --remove on a scale-to-zero service with cap 3:

  • reads in the help that the cap is kept;
  • then gets the new success line (src/commands/compute.ts:679) saying replica cap is now 2;
  • and on the next deploy, the plane's cap really does drop to 2.

The help and the output contradict each other, and the help is the one that's wrong.

The fix is instacloud-platform PR 642 ("removing an instance writes back the plane's replica cap"). It's still open and unreviewed. Once it ships, the help sentence and the printed cap both become true. Either merge this after 642 is deployed, or drop the "a scale-to-zero one keeps its cap" clause until it is. The always-on half ("runs one fewer") is correct today: the plane sets both cap and count to old-1 for a fixed service.

Evidence

read-the-code — src/index.ts:398-400, src/commands/compute.ts:669-680; instacloud-platform (origin/main) src/provisioning/services.ts:1677-1705, src/adapters/insta-compute.ts:1742-1743; instacloud-compute internal/executor/executor.go:1365-1378, internal/meta/meta.go:202-223 (SetReplicas keeps the cap; Autoscaled = not always-on, no volume, routed, cap > 1); gh pr view 642 -R InsForge/instacloud-platform (OPEN, REVIEW_REQUIRED) and its diff; instacloud-compute PR 407 merged 2026-10-07

@Fermionic-Lyu

Copy link
Copy Markdown
Member Author

InsForge/instacloud-platform#642 is merged and deployed to prod (62b25c62, run 37709825443). machine_count now follows the plane's cap after --remove, so the help sentence and the printed cap both hold as written. No code change since 9993349.

@agent-zhang-beihai agent-zhang-beihai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed by Wang Miao

I'm approving: the remaining finding was resolved on the platform side. This PR rewords insta compute scale around a replica cap, and --remove now explains what happens to the cap and prints it. Last round I held it because the help said removing an instance from a scale-to-zero service keeps its cap, while the platform still lowered it by one. instacloud-platform PR 642 has since merged (62b25c6), and its removeInstance now writes back the cap the compute plane reports (machineCount: replicaCap ?? current - 1). The author says that is deployed to prod. I couldn't open the deploy run, but the merged code is what makes the help text and the printed cap agree. The author was right that the CLI wording was correct once that platform fix landed.

With fresh eyes I checked what I hadn't last time:

  • Workers: a port-0 worker is always-on (instacloud-platform src/provisioning/services.ts:60-65, CLI --port 0 help), so "scale-to-zero keeps its cap" never reaches a worker, which the plane treats as fixed.
  • Skills doc: the agent-facing insta/cli-reference.md in instacloud-skills already describes the cap the same way.
  • Old message strings: nothing in this repo, instacloud-e2e, instacloud-agent-e2e or instacloud-mcp matches the removed messages.
  • Tests: I couldn't run test/compute-scale-remove.test.ts because this checkout has no node_modules.

No findings.

@jwfing jwfing left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM - approved.

@Fermionic-Lyu
Fermionic-Lyu merged commit 3ac5be7 into main Oct 8, 2026
4 of 5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants