Skip to content

DX-596-together-dedicated-endpoints: sync with mintlify-docs#991 - #55

Closed
zainhas wants to merge 1 commit into
mainfrom
docs-sync/together-dedicated-endpoints/mintlify-docs-pr-991
Closed

zainhas wants to merge 1 commit into
mainfrom
docs-sync/together-dedicated-endpoints/mintlify-docs-pr-991

Conversation

@zainhas

@zainhas zainhas commented Jul 16, 2026

Copy link
Copy Markdown
Collaborator

Sync the together-dedicated-endpoints skill with the DE 2.0 (dedicated model inference) documentation rewrite merged in mintlify-docs#991.

Changed docs files that triggered this sync

  • docs/dedicated-endpoints/overview.mdx — rewritten for DMI 2.0
  • docs/dedicated-endpoints/quickstart.mdx — CLI-first tg beta endpoints deploy walkthrough
  • docs/dedicated-endpoints/manage.mdx — new endpoint/deployment lifecycle
  • docs/dedicated-endpoints/scaling.mdx — new autoscaling metrics and windows
  • docs/dedicated-endpoints/concepts.mdx (new) — resource model, cold starts, sticky routing
  • docs/dedicated-endpoints/configs.mdx (new) — deployment profiles and configs
  • docs/dedicated-endpoints/pricing.mdx (new) — DMI billing and instance-type table
  • docs/dedicated-endpoints/requests.mdx (new) — inference at api-inference.together.ai/v1
  • docs/dedicated-endpoints/route-traffic.mdx (new) — routing pipeline + stickiness
  • docs/dedicated-endpoints/split-traffic.mdx (new) — weighted traffic splits
  • docs/dedicated-endpoints/ab-tests.mdx (new) — A/B experiments (abx_)
  • docs/dedicated-endpoints/shadow-experiments.mdx (new) — shadow experiments (exp_)
  • docs/dedicated-endpoints/rollouts.mdx (new) — metric-gated rollouts (rol_)
  • docs/dedicated-endpoints/monitoring.mdx (new) — events feed + Prometheus metrics endpoint
  • docs/dedicated-endpoints/custom-models.mdx — v2 fine-tuned model upload flow
  • docs/dedicated-endpoints/adapter.mdx — v2 LoRA adapter upload flow
  • docs/dedicated-endpoints/models.mdx — new supported-models catalog with deployment profiles
  • docs/dedicated-endpoints/migrate-from-v1.mdx (new) — v1 → v2 migration guide
  • docs/dedicated-endpoints/v1/* (new) — v1 legacy pages retained under /v1/

Skill changes

  • SKILL.md rewritten to make v2 (DMI) the default while noting v1 remains supported through the end of 2026. Quick-routing menu adds the v2 CLI workflows (deploy, split, A/B, shadow, rollout, custom-model + LoRA upload). High-signal rules updated for CLI 2.24.0+, project scope, ep_/dep_ IDs, min:0 requires max:0, autoscaling metric units, canary-only metric gates, and prompt caching being on by default.
  • references/api-reference.md restructured v2-first: DMI resource model, full tg beta surface, deploy flags, deployment states, autoscaling metrics + rate limits + stabilization windows, configs and profiles, instance types + headroom, model/adapter uploads, observability (events feed + Prometheus-compatible metrics endpoint), smart-delete rules by ID prefix, and a concise v1 legacy section with the command mapping table.
  • references/traffic-routing.md (new) documents the full v2 routing surface: how routing works (traffic split → A/B re-sample → route), stickiness via prompt_cache_key, basic weights and split-traffic semantics, A/B tests with ramp/promote/delete, shadow experiments (all four sampling strategies), and rollouts with canary/blue-green/rolling strategies, metric gates, per-replica normalization for inflight_requests, pause categories, state machine, and roll-back guidance.
  • references/models-and-configs.md (new) covers architecture / weight / config layering, tg beta models public and tg beta models configs responses, the v2 fine-tuned model upload walkthrough (create → upload → poll → deploy), LoRA adapter uploads with the --type adapter flag, the doubled-slug naming pitfall, and speculative-decoding as a config property.
  • references/hardware-options.md adds the new v2 instance-type IDs (1xnvidia-h100-80gb, 1xnvidia-h200-141gb, 1xnvidia-b200-180gb, 1xnvidia-gb300-280gb, 1xnvidia-b300-280gb), refreshes the pricing table (H100 $5.49/hr, B200 $8.99/hr, others contact-sales), adds the DMI vs serverless break-even note, and keeps the v1 underscore hardware IDs / v1 SDK query examples for legacy users.
  • references/dedicated-models.md adds tg beta models public --product DEDICATED as the primary listing method, preserves the current model snapshot (chat, image, transcription, moderation, rerank), and cross-references the v2 upload workflow.
  • scripts/deploy_v2.sh (new) is an end-to-end v2 CLI walkthrough: tg beta endpoints deploy, poll for DEPLOYMENT_STATE_READY, send a chat completion to api-inference.together.ai/v1, tear down with tg beta endpoints rm --force.
  • scripts/manage_endpoint.py, manage_endpoint.ts, deploy_finetuned.py, upload_custom_model.py docstrings marked as v1 (legacy) helpers and pointing users to the v2 CLI equivalents.

Notes for the reviewer

  • The v2 SDK surface (client.beta.endpoints.*, client.beta.models.*) is still being published; the skill leads with the CLI (tg beta) and raw HTTP for automation, and calls this out explicitly in high-signal rules.
  • Multi-LoRA serving has moved to the v1 section per the upstream PR; v2 supports single-adapter deployment via tg beta models create --type adapter. Both are covered.
  • A separate manual effort (skills#50) proposes a new together-dedicated-model-inference directory; this sync updates the existing mapped together-dedicated-endpoints skill in place per the sync map and hard rules (no new directories).
  • Two other DE PRs remain open against v1-shaped content: skills#44 (multi-LoRA hot-swap, now v1-only per upstream) and skills#51 (LoRA-enabled base models table). Neither overlaps with the files touched here.

Generated by the Sync Skills Cursor Automation. Please review before merging.

Sync the together-dedicated-endpoints skill with the DE 2.0 (dedicated
model inference) documentation rewrite merged in mintlify-docs#991.

- SKILL.md: rewritten to make v2 (DMI) the default while noting v1 remains
  supported through end-of-2026. Adds new capabilities (traffic splits,
  A/B tests, shadow experiments, metric-gated rollouts, model + LoRA
  uploads via tg beta) to the quick-routing menu, and refreshes the
  high-signal rules (CLI 2.24.0+, project scope, ep_/dep_ IDs, min:0
  requires max:0, autoscaling metric units, canary-only metric gates,
  prompt caching default-on).

- references/api-reference.md: restructured as v2-first, with the DMI
  resource model (project/model/config/endpoint/deployment/replica),
  full tg beta CLI surface, deploy flags, deployment states, autoscaling
  metrics + timing (rate limits, stabilization windows), configs and
  profiles, instance types + headroom, model/adapter upload flow,
  observability (events + Prometheus-compatible metrics endpoint), smart
  delete rules, and a concise legacy v1 API section with the command
  mapping table.

- references/traffic-routing.md (new): documents the entire v2 routing
  surface — how routing works (traffic split -> A/B re-sample -> route),
  stickiness / prompt_cache_key, basic weights + split-traffic
  semantics, A/B tests (abx_) with ramp/promote/delete, shadow
  experiments (exp_) with all four sampling strategies, and rollouts
  (rol_) with strategies, metric gates, per-replica normalization for
  inflight_requests, pause categories, state machine, and roll-back
  guidance.

- references/models-and-configs.md (new): architecture / weight / config
  layering and deployment-profile concept, tg beta models public and
  configs commands with full response shapes, v2 fine-tuned model upload
  walkthrough (create -> upload / remote-uploads -> poll -> deploy),
  LoRA adapter uploads (--type adapter, doubled-slug pitfall, multi-LoRA
  note), profile selection guidance, and speculative-decoding as a
  config property.

- references/hardware-options.md: adds the new v2 instance-type IDs
  (1xnvidia-h100-80gb, 1xnvidia-h200-141gb, 1xnvidia-b200-180gb,
  1xnvidia-gb300-280gb, 1xnvidia-b300-280gb), the DMI pricing table
  ($5.49/hr H100, $8.99/hr B200, others contact-sales), the DMI vs
  serverless break-even note, headroom explanation, and keeps v1
  underscore hardware IDs / v1 SDK query examples for legacy users.

- references/dedicated-models.md: adds tg beta models public --product
  DEDICATED as the primary listing method, preserves the current model
  snapshot (chat, image, transcription, moderation, rerank), and
  cross-references the v2 upload workflow.

- scripts/deploy_v2.sh (new): end-to-end v2 CLI walkthrough — tg beta
  endpoints deploy, poll for DEPLOYMENT_STATE_READY on the deployment,
  send a chat completion to api-inference.together.ai/v1, and tear the
  endpoint down with tg beta endpoints rm --force.

- scripts/manage_endpoint.py, manage_endpoint.ts, deploy_finetuned.py,
  upload_custom_model.py: docstrings updated to mark them as v1 (legacy)
  helpers and point users to the v2 CLI equivalents in
  scripts/deploy_v2.sh and references/models-and-configs.md.

Generated by the Sync Skills Cursor Automation.
Please review before merging.

Refs: togethercomputer/mintlify-docs#991
Co-authored-by: Mo King <mking@together.ai>
@zainhas zainhas self-assigned this Jul 16, 2026
@zainhas
zainhas requested a review from muhsinking July 16, 2026 00:28
@broly-code-security-scanner

Copy link
Copy Markdown

Broly Security Scan

Note

✅ Clean scan
No vulnerabilities detected in this PR.

Note

Re-scan this PR anytime with /broly scan — useful after /broly undismiss, or to refresh findings without a new push.

Broly — SAST (zai-org/GLM-5.2) · Secrets · SCA · GH Actions (zizmor) · Containers · SBOM · Powered by Together AI

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Heads-up from Sync Skills automation: mintlify-docs#1244 ("docs: sync DE 2.0 Python SDK release from together-py#448") merged into main on 2026-07-16 with updates that overlap this v2-rewrite PR. Before merging skills#55, please fold the following in (or leave a follow-up commit).

Updates from mintlify-docs#1244:

  1. client.beta.endpoints.* and client.beta.models.* are now shipped in the published together Python SDK. Only client.beta.endpoints.rollouts.* remains unpublished (kept behind a validator-ignore in the docs).
  2. Python SDK project scoping — pass project_id to Together() or set TOGETHER_PROJECT_ID. Otherwise call client.whoami().project_id before project-scoped API calls.
  3. New global --project [string] CLI flag documented under CLI global parameters (fallback to TOGETHER_PROJECT_ID; without either, read-only commands use the API key's default project, mutating commands may prompt for confirmation or require an explicit project in --json/--non-interactive mode).
  4. tg whoami --json now includes a user_id field (present for user-account API keys, omitted for service or organization-default keys).
Open in Web View Automation 

Sent by Cursor Automation: Sync mintlify-docs to skills

- Autoscaling, auto-shutdown, prompt caching, and speculative decoding materially affect operations and cost.
- For custom or fine-tuned models, do not skip the intermediate verification steps before deployment.
- **v2 CLI requires Together CLI 2.24.0+**. Install with `uv tool install "together[cli]"` (or `pip install together`). Every v2 management command lives under `tg beta`.
- **v2 SDK surface (`client.beta.endpoints.*`, `client.beta.models.*`) is still being published**. Prefer the CLI or raw HTTP for v2 automation until the SDK stabilizes; the v1 SDK (`client.endpoints.*`) continues to work only against v1 endpoints.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale after mintlify-docs#1244: client.beta.endpoints.* and client.beta.models.* are now in the published together Python SDK. Only rollouts (client.beta.endpoints.rollouts.*) are still unpublished.

Suggested rewrite:

v2 SDK is available for client.beta.endpoints.* and client.beta.models.* in the published together Python SDK (together>=2.x). Rollouts (client.beta.endpoints.rollouts.*) are the one v2 surface still awaiting SDK release — use the CLI or raw HTTP for rollout automation.

- For custom or fine-tuned models, do not skip the intermediate verification steps before deployment.
- **v2 CLI requires Together CLI 2.24.0+**. Install with `uv tool install "together[cli]"` (or `pip install together`). Every v2 management command lives under `tg beta`.
- **v2 SDK surface (`client.beta.endpoints.*`, `client.beta.models.*`) is still being published**. Prefer the CLI or raw HTTP for v2 automation until the SDK stabilizes; the v1 SDK (`client.endpoints.*`) continues to work only against v1 endpoints.
- **Project scope matters**. `tg beta` and the v2 management API read the project from `TOGETHER_PROJECT_ID` or the `--project` flag; without either, the CLI falls back to the project associated with your API key and prompts for confirmation in interactive shells.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

mintlify-docs#1244 also documents the Python SDK project scoping (this rule currently only covers CLI + API):

The Python SDK also reads the project from TOGETHER_PROJECT_ID or an explicit project_id= kwarg on Together(). When neither is set and a call needs a project, resolve one first with client.whoami().project_id.

Consider extending this bullet to mention project_id= on the SDK client and the client.whoami().project_id fallback.

```

All v2 management commands live under `tg beta`. The Python SDK is imported as `together` (v2 SDK
uses `client.beta.endpoints.*` and `client.beta.models.*` namespaces; some surfaces are still being

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale after mintlify-docs#1244: beta.endpoints.* and beta.models.* now ship in the published together SDK. Narrow the caveat to rollouts only.

Suggested rewrite:

The Python SDK is imported as together — the v2 surface uses client.beta.endpoints.* and client.beta.models.* namespaces (now published in together>=2.x). client.beta.endpoints.rollouts.* is the one v2 surface still awaiting SDK release; use the CLI or raw HTTP for rollout automation.

```bash
export TOGETHER_PROJECT_ID=proj_abc123 # explicit
tg beta endpoints ls --project proj_abc123 # per-command override
tg whoami # inspect the resolved project

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

mintlify-docs#1244 documents the Python SDK project scoping too. Consider adding a Python example alongside the shell block:

from together import Together

client = Together(project_id="proj_abc123")   # or set TOGETHER_PROJECT_ID
# Otherwise resolve one first:
project_id = client.whoami().project_id

Also worth noting under tg whoami: the JSON response now includes a user_id field (present for user-account API keys, omitted for service or organization-default keys).

@zainhas zainhas closed this Jul 16, 2026
@zainhas
zainhas deleted the docs-sync/together-dedicated-endpoints/mintlify-docs-pr-991 branch July 16, 2026 17:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant