Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 8 additions & 4 deletions docs/runware_serverless_apps_scale.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,10 +6,14 @@ Scale a serverless application

Patch live worker configuration for a serverless application.

Omitted flags are left unchanged. Configuration changes take effect on the
next scaler cycle; this command does not wait for a rollout.
Omitted flags are left unchanged. A configuration change records a new version
with the same image and rolls the workload when the app is active or
initializing. A failed app is moved to initializing and rolled. A stopped or
stopping app applies the change on resume. This command does not wait for that
rollout. A change during an in-flight rollout returns 409.

The server rejects unsupported or invalid fields with HTTP 422.
The server rejects unsupported or invalid fields with HTTP 422. A capacity
increase without enough credit returns 402.

```
runware serverless apps scale <appId> [flags]
Expand All @@ -24,7 +28,7 @@ runware serverless apps scale <appId> [flags]
# scale to zero and raise idle TTL
runware serverless apps scale my-app --min-workers 0 --idle-ttl 120

# change GPU type (applies to newly created workers)
# change GPU type; the rollout replaces workers with the new type
runware serverless apps scale my-app --gpu-type h100
```

Expand Down
70 changes: 43 additions & 27 deletions docs/runware_serverless_deploy.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,10 @@ Create or update a serverless application from Python code or a container source
A first deploy with a new --id creates the application. A later deploy with the
same --id uploads a new source, records version N+1, and rolls it when the
build is ready. Create-only flags (--gpu-type, worker settings, --volume,
--env, --env-file, --name) apply only to create; passing them when the
application already exists is an error. Change workers with 'apps scale' and
environment with 'apps env'. A source update on a stopped application is 409.
--secret, --env, --env-file, --name) apply only to create; passing them when
the application already exists is an error. Change workers with 'apps scale',
attach a secret later with 'secrets attach', and change environment with
'apps env'. A source update on a stopped application is 409.

A code deploy takes a Python entry file. The whole source directory is zipped
and submitted as the application source, so the entry file can import its own
Expand Down Expand Up @@ -49,13 +50,20 @@ returns 409 and does not store the value. Prefer --env-file for anything
secret: a value passed as --env is visible in the process list and recorded
in shell history.

Anything the app downloads at runtime belongs on a --volume. The app runs in a
sandbox whose filesystem is part of the checkpointed state, so an unmounted
download is copied into every checkpoint and fetched again on every cold start.
A volume keeps it out of both.
Anything the app downloads at runtime belongs on a --volume. Volumes are set
at create and cannot be changed afterwards: a later deploy that passes
--volume is rejected. The app runs in a sandbox whose filesystem is part of
the checkpointed state, so an unmounted download is copied into every
checkpoint and fetched again on every cold start. A volume keeps it out of
both.

Worker settings are supplied via flags on create. Endpoints are derived
server-side from the SDK (code) or from container.yaml (container).
--secret NAME, or NAME=ENV_VAR, attaches an existing organisation secret at
create so the first rollout carries it. Repeat the flag for more than one.
Attach or detach later with 'secrets attach' and 'secrets detach'.

Worker settings are supplied via flags on create, including a fallback GPU
type and the idle-worker buffer. Endpoints are derived server-side from the
SDK (code) or from container.yaml (container).

A code app's endpoint path is its handler's method name with underscores turned
into hyphens, so renaming a method moves a public endpoint and 404s its callers.
Expand Down Expand Up @@ -91,6 +99,10 @@ runware serverless deploy [file] [flags]
runware serverless deploy model.py --id my-app --gpu-type l40s \
--volume /root/.cache/huggingface

# attach an existing secret and keep one idle worker warm
runware serverless deploy ./app.py --id my-app --gpu-type h100 \
--secret HF_TOKEN=HUGGING_FACE_HUB_TOKEN --min-available-workers 1

# override worker settings and base image
runware serverless deploy ./app.py --id my-app --name "My App" \
--max-workers 2 --idle-ttl 120 --gpu-type h100 \
Expand All @@ -106,24 +118,28 @@ runware serverless deploy [file] [flags]
### Options

```
--base-image string Builder base image (code deploys only; needs Python 3.12 or newer) (default "python:3.12-slim")
--container string Directory whose root contains Dockerfile and container.yaml
--env stringArray Environment variable as KEY=VALUE (repeatable)
--env-file stringArray File of KEY=VALUE lines to read environment variables from (repeatable)
--gpu-type string GPU type ID (see 'serverless gpus'; required when creating)
--gpus-per-worker int32 GPUs allocated per worker (1, 2, 4, or 8) (default 1)
-h, --help help for deploy
--id string Application ID (immutable, lowercase slug)
--idle-ttl int32 Idle TTL in seconds before scaling down (default 60)
--max-workers int32 Maximum number of workers (default 1)
--min-workers int32 Minimum number of workers
--name string Display name (defaults to --id)
--poll-interval duration Polling interval when waiting for the application (default 2s)
--requirement stringArray Additional pip package to install (repeatable; code deploys only)
--scaling-delay int32 Scaling delay in seconds (default 10)
--src-dir string Directory to package as the application source (default: the working directory; code deploys only)
--volume stringArray Absolute path inside the app backed by persistent node-local storage (repeatable)
--wait Poll until the application is active or failed
--available-workers-pct int32 Idle-worker buffer as a percentage of load (0-100)
--base-image string Builder base image (code deploys only; needs Python 3.12 or newer) (default "python:3.12-slim")
--container string Directory whose root contains Dockerfile and container.yaml
--env stringArray Environment variable as KEY=VALUE (repeatable)
--env-file stringArray File of KEY=VALUE lines to read environment variables from (repeatable)
--fallback-gpu-type string Secondary GPU type if the preferred type is unavailable
--gpu-type string GPU type ID (see 'serverless gpus'; required when creating)
--gpus-per-worker int32 GPUs allocated per worker (1, 2, 4, or 8) (default 1)
-h, --help help for deploy
--id string Application ID (immutable, lowercase slug)
--idle-ttl int32 Idle TTL in seconds before scaling down (default 60)
--max-workers int32 Maximum number of workers (default 1)
--min-available-workers int32 Minimum idle workers kept as a buffer
--min-workers int32 Minimum number of workers
--name string Display name (defaults to --id)
--poll-interval duration Polling interval when waiting for the application (default 2s)
--requirement stringArray Additional pip package to install (repeatable; code deploys only)
--scaling-delay int32 Scaling delay in seconds (default 10)
--secret stringArray Organisation secret to attach at create, as NAME or NAME=ENV_VAR (repeatable)
--src-dir string Directory to package as the application source (default: the working directory; code deploys only)
--volume stringArray Absolute path inside the app backed by persistent node-local storage; immutable after create (repeatable)
--wait Poll until the application is active or failed
```

### Options inherited from parent commands
Expand Down
12 changes: 8 additions & 4 deletions internal/cmd/serverless/apps_scale.go
Original file line number Diff line number Diff line change
Expand Up @@ -34,17 +34,21 @@ func newAppsScaleCmd(logger *log.Logger) *cobra.Command {
Short: "Scale a serverless application",
Long: `Patch live worker configuration for a serverless application.

Omitted flags are left unchanged. Configuration changes take effect on the
next scaler cycle; this command does not wait for a rollout.
Omitted flags are left unchanged. A configuration change records a new version
with the same image and rolls the workload when the app is active or
initializing. A failed app is moved to initializing and rolled. A stopped or
stopping app applies the change on resume. This command does not wait for that
rollout. A change during an in-flight rollout returns 409.

The server rejects unsupported or invalid fields with HTTP 422.`,
The server rejects unsupported or invalid fields with HTTP 422. A capacity
increase without enough credit returns 402.`,
Example: ` # set the worker cap
runware serverless apps scale my-app --max-workers 2

# scale to zero and raise idle TTL
runware serverless apps scale my-app --min-workers 0 --idle-ttl 120

# change GPU type (applies to newly created workers)
# change GPU type; the rollout replaces workers with the new type
runware serverless apps scale my-app --gpu-type h100`,
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
Expand Down
2 changes: 1 addition & 1 deletion internal/cmd/serverless/apps_scale_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ func TestWorkerConfigPatchFromFlags_EachFlag(t *testing.T) {
{[]string{"--min-workers", "0"}, "minWorkers", float64(0)},
{[]string{"--idle-ttl", "120"}, "idleTtlSecs", float64(120)},
{[]string{"--scaling-delay", "15"}, "scalingDelaySecs", float64(15)},
{[]string{"--gpu-type", testGPUType}, "gpuType", testGPUType},
{[]string{testGPUTypeFlag, testGPUType}, "gpuType", testGPUType},
{[]string{"--gpus-per-worker", "2"}, "gpusPerWorker", float64(2)},
{[]string{"--fallback-gpu-type", testGPUType}, "fallbackGpuType", testGPUType},
{[]string{"--min-available-workers", "1"}, "minAvailableWorkers", float64(1)},
Expand Down
Loading
Loading