Repository navigation
Conversation
Sync the together-dedicated-endpoints skill with the DE 2.0 (dedicated model inference) documentation rewrite merged in mintlify-docs#991. - SKILL.md: rewritten to make v2 (DMI) the default while noting v1 remains supported through end-of-2026. Adds new capabilities (traffic splits, A/B tests, shadow experiments, metric-gated rollouts, model + LoRA uploads via tg beta) to the quick-routing menu, and refreshes the high-signal rules (CLI 2.24.0+, project scope, ep_/dep_ IDs, min:0 requires max:0, autoscaling metric units, canary-only metric gates, prompt caching default-on). - references/api-reference.md: restructured as v2-first, with the DMI resource model (project/model/config/endpoint/deployment/replica), full tg beta CLI surface, deploy flags, deployment states, autoscaling metrics + timing (rate limits, stabilization windows), configs and profiles, instance types + headroom, model/adapter upload flow, observability (events + Prometheus-compatible metrics endpoint), smart delete rules, and a concise legacy v1 API section with the command mapping table. - references/traffic-routing.md (new): documents the entire v2 routing surface — how routing works (traffic split -> A/B re-sample -> route), stickiness / prompt_cache_key, basic weights + split-traffic semantics, A/B tests (abx_) with ramp/promote/delete, shadow experiments (exp_) with all four sampling strategies, and rollouts (rol_) with strategies, metric gates, per-replica normalization for inflight_requests, pause categories, state machine, and roll-back guidance. - references/models-and-configs.md (new): architecture / weight / config layering and deployment-profile concept, tg beta models public and configs commands with full response shapes, v2 fine-tuned model upload walkthrough (create -> upload / remote-uploads -> poll -> deploy), LoRA adapter uploads (--type adapter, doubled-slug pitfall, multi-LoRA note), profile selection guidance, and speculative-decoding as a config property. - references/hardware-options.md: adds the new v2 instance-type IDs (1xnvidia-h100-80gb, 1xnvidia-h200-141gb, 1xnvidia-b200-180gb, 1xnvidia-gb300-280gb, 1xnvidia-b300-280gb), the DMI pricing table ($5.49/hr H100, $8.99/hr B200, others contact-sales), the DMI vs serverless break-even note, headroom explanation, and keeps v1 underscore hardware IDs / v1 SDK query examples for legacy users. - references/dedicated-models.md: adds tg beta models public --product DEDICATED as the primary listing method, preserves the current model snapshot (chat, image, transcription, moderation, rerank), and cross-references the v2 upload workflow. - scripts/deploy_v2.sh (new): end-to-end v2 CLI walkthrough — tg beta endpoints deploy, poll for DEPLOYMENT_STATE_READY on the deployment, send a chat completion to api-inference.together.ai/v1, and tear the endpoint down with tg beta endpoints rm --force. - scripts/manage_endpoint.py, manage_endpoint.ts, deploy_finetuned.py, upload_custom_model.py: docstrings updated to mark them as v1 (legacy) helpers and point users to the v2 CLI equivalents in scripts/deploy_v2.sh and references/models-and-configs.md. Generated by the Sync Skills Cursor Automation. Please review before merging. Refs: togethercomputer/mintlify-docs#991 Co-authored-by: Mo King <mking@together.ai>
Broly Security ScanNote ✅ Clean scan Note Re-scan this PR anytime with
|
There was a problem hiding this comment.
Heads-up from Sync Skills automation: mintlify-docs#1244 ("docs: sync DE 2.0 Python SDK release from together-py#448") merged into main on 2026-07-16 with updates that overlap this v2-rewrite PR. Before merging skills#55, please fold the following in (or leave a follow-up commit).
Updates from mintlify-docs#1244:
client.beta.endpoints.*andclient.beta.models.*are now shipped in the publishedtogetherPython SDK. Onlyclient.beta.endpoints.rollouts.*remains unpublished (kept behind a validator-ignore in the docs).- Python SDK project scoping — pass
project_idtoTogether()or setTOGETHER_PROJECT_ID. Otherwise callclient.whoami().project_idbefore project-scoped API calls. - New global
--project [string]CLI flag documented under CLI global parameters (fallback toTOGETHER_PROJECT_ID; without either, read-only commands use the API key's default project, mutating commands may prompt for confirmation or require an explicit project in--json/--non-interactivemode). tg whoami --jsonnow includes auser_idfield (present for user-account API keys, omitted for service or organization-default keys).
Sent by Cursor Automation: Sync mintlify-docs to skills
| - Autoscaling, auto-shutdown, prompt caching, and speculative decoding materially affect operations and cost. | ||
| - For custom or fine-tuned models, do not skip the intermediate verification steps before deployment. | ||
| - **v2 CLI requires Together CLI 2.24.0+**. Install with `uv tool install "together[cli]"` (or `pip install together`). Every v2 management command lives under `tg beta`. | ||
| - **v2 SDK surface (`client.beta.endpoints.*`, `client.beta.models.*`) is still being published**. Prefer the CLI or raw HTTP for v2 automation until the SDK stabilizes; the v1 SDK (`client.endpoints.*`) continues to work only against v1 endpoints. |
There was a problem hiding this comment.
Stale after mintlify-docs#1244: client.beta.endpoints.* and client.beta.models.* are now in the published together Python SDK. Only rollouts (client.beta.endpoints.rollouts.*) are still unpublished.
Suggested rewrite:
v2 SDK is available for
client.beta.endpoints.*andclient.beta.models.*in the publishedtogetherPython SDK (together>=2.x). Rollouts (client.beta.endpoints.rollouts.*) are the one v2 surface still awaiting SDK release — use the CLI or raw HTTP for rollout automation.
| - For custom or fine-tuned models, do not skip the intermediate verification steps before deployment. | ||
| - **v2 CLI requires Together CLI 2.24.0+**. Install with `uv tool install "together[cli]"` (or `pip install together`). Every v2 management command lives under `tg beta`. | ||
| - **v2 SDK surface (`client.beta.endpoints.*`, `client.beta.models.*`) is still being published**. Prefer the CLI or raw HTTP for v2 automation until the SDK stabilizes; the v1 SDK (`client.endpoints.*`) continues to work only against v1 endpoints. | ||
| - **Project scope matters**. `tg beta` and the v2 management API read the project from `TOGETHER_PROJECT_ID` or the `--project` flag; without either, the CLI falls back to the project associated with your API key and prompts for confirmation in interactive shells. |
There was a problem hiding this comment.
mintlify-docs#1244 also documents the Python SDK project scoping (this rule currently only covers CLI + API):
The Python SDK also reads the project from
TOGETHER_PROJECT_IDor an explicitproject_id=kwarg onTogether(). When neither is set and a call needs a project, resolve one first withclient.whoami().project_id.
Consider extending this bullet to mention project_id= on the SDK client and the client.whoami().project_id fallback.
| ``` | ||
|
|
||
| All v2 management commands live under `tg beta`. The Python SDK is imported as `together` (v2 SDK | ||
| uses `client.beta.endpoints.*` and `client.beta.models.*` namespaces; some surfaces are still being |
There was a problem hiding this comment.
Stale after mintlify-docs#1244: beta.endpoints.* and beta.models.* now ship in the published together SDK. Narrow the caveat to rollouts only.
Suggested rewrite:
The Python SDK is imported as
together— the v2 surface usesclient.beta.endpoints.*andclient.beta.models.*namespaces (now published intogether>=2.x).client.beta.endpoints.rollouts.*is the one v2 surface still awaiting SDK release; use the CLI or raw HTTP for rollout automation.
| ```bash | ||
| export TOGETHER_PROJECT_ID=proj_abc123 # explicit | ||
| tg beta endpoints ls --project proj_abc123 # per-command override | ||
| tg whoami # inspect the resolved project |
There was a problem hiding this comment.
mintlify-docs#1244 documents the Python SDK project scoping too. Consider adding a Python example alongside the shell block:
from together import Together
client = Together(project_id="proj_abc123") # or set TOGETHER_PROJECT_ID
# Otherwise resolve one first:
project_id = client.whoami().project_idAlso worth noting under tg whoami: the JSON response now includes a user_id field (present for user-account API keys, omitted for service or organization-default keys).


Sync the
together-dedicated-endpointsskill with the DE 2.0 (dedicated model inference) documentation rewrite merged in mintlify-docs#991.Changed docs files that triggered this sync
docs/dedicated-endpoints/overview.mdx— rewritten for DMI 2.0docs/dedicated-endpoints/quickstart.mdx— CLI-firsttg beta endpoints deploywalkthroughdocs/dedicated-endpoints/manage.mdx— new endpoint/deployment lifecycledocs/dedicated-endpoints/scaling.mdx— new autoscaling metrics and windowsdocs/dedicated-endpoints/concepts.mdx(new) — resource model, cold starts, sticky routingdocs/dedicated-endpoints/configs.mdx(new) — deployment profiles and configsdocs/dedicated-endpoints/pricing.mdx(new) — DMI billing and instance-type tabledocs/dedicated-endpoints/requests.mdx(new) — inference atapi-inference.together.ai/v1docs/dedicated-endpoints/route-traffic.mdx(new) — routing pipeline + stickinessdocs/dedicated-endpoints/split-traffic.mdx(new) — weighted traffic splitsdocs/dedicated-endpoints/ab-tests.mdx(new) — A/B experiments (abx_)docs/dedicated-endpoints/shadow-experiments.mdx(new) — shadow experiments (exp_)docs/dedicated-endpoints/rollouts.mdx(new) — metric-gated rollouts (rol_)docs/dedicated-endpoints/monitoring.mdx(new) — events feed + Prometheus metrics endpointdocs/dedicated-endpoints/custom-models.mdx— v2 fine-tuned model upload flowdocs/dedicated-endpoints/adapter.mdx— v2 LoRA adapter upload flowdocs/dedicated-endpoints/models.mdx— new supported-models catalog with deployment profilesdocs/dedicated-endpoints/migrate-from-v1.mdx(new) — v1 → v2 migration guidedocs/dedicated-endpoints/v1/*(new) — v1 legacy pages retained under/v1/Skill changes
SKILL.mdrewritten to make v2 (DMI) the default while noting v1 remains supported through the end of 2026. Quick-routing menu adds the v2 CLI workflows (deploy, split, A/B, shadow, rollout, custom-model + LoRA upload). High-signal rules updated for CLI 2.24.0+, project scope,ep_/dep_IDs,min:0requiresmax:0, autoscaling metric units, canary-only metric gates, and prompt caching being on by default.references/api-reference.mdrestructured v2-first: DMI resource model, fulltg betasurface, deploy flags, deployment states, autoscaling metrics + rate limits + stabilization windows, configs and profiles, instance types + headroom, model/adapter uploads, observability (events feed + Prometheus-compatible metrics endpoint), smart-delete rules by ID prefix, and a concise v1 legacy section with the command mapping table.references/traffic-routing.md(new) documents the full v2 routing surface: how routing works (traffic split → A/B re-sample → route), stickiness viaprompt_cache_key, basic weights and split-traffic semantics, A/B tests with ramp/promote/delete, shadow experiments (all four sampling strategies), and rollouts with canary/blue-green/rolling strategies, metric gates, per-replica normalization forinflight_requests, pause categories, state machine, and roll-back guidance.references/models-and-configs.md(new) covers architecture / weight / config layering,tg beta models publicandtg beta models configsresponses, the v2 fine-tuned model upload walkthrough (create → upload → poll → deploy), LoRA adapter uploads with the--type adapterflag, the doubled-slug naming pitfall, and speculative-decoding as a config property.references/hardware-options.mdadds the new v2 instance-type IDs (1xnvidia-h100-80gb,1xnvidia-h200-141gb,1xnvidia-b200-180gb,1xnvidia-gb300-280gb,1xnvidia-b300-280gb), refreshes the pricing table (H100 $5.49/hr, B200 $8.99/hr, others contact-sales), adds the DMI vs serverless break-even note, and keeps the v1 underscore hardware IDs / v1 SDK query examples for legacy users.references/dedicated-models.mdaddstg beta models public --product DEDICATEDas the primary listing method, preserves the current model snapshot (chat, image, transcription, moderation, rerank), and cross-references the v2 upload workflow.scripts/deploy_v2.sh(new) is an end-to-end v2 CLI walkthrough:tg beta endpoints deploy, poll forDEPLOYMENT_STATE_READY, send a chat completion toapi-inference.together.ai/v1, tear down withtg beta endpoints rm --force.scripts/manage_endpoint.py,manage_endpoint.ts,deploy_finetuned.py,upload_custom_model.pydocstrings marked as v1 (legacy) helpers and pointing users to the v2 CLI equivalents.Notes for the reviewer
client.beta.endpoints.*,client.beta.models.*) is still being published; the skill leads with the CLI (tg beta) and raw HTTP for automation, and calls this out explicitly in high-signal rules.tg beta models create --type adapter. Both are covered.together-dedicated-model-inferencedirectory; this sync updates the existing mappedtogether-dedicated-endpointsskill in place per the sync map and hard rules (no new directories).Generated by the Sync Skills Cursor Automation. Please review before merging.