amplihack-remote API¶
The amplihack-remote crate is the native Rust API behind
amplihack remote. It owns remote execution behavior, VM pooling, detached
session state, output capture, and result integration. CLI code in
crates/amplihack-cli only parses arguments, calls this crate, and renders the
returned values.
Contents¶
- Design contract
- Target implementation contract
- High-level command API
- Synchronous execution
- Detached sessions
- VM pooling
- Idle VM cleanup
- State management
- Errors
- Testing seams
Design contract¶
amplihack-remote exposes public functions for the six supported remote
commands:
execliststartoutputkillstatus
The crate also keeps lower-level public building blocks for consumers that need direct control over packaging, provisioning, execution, integration, or session state.
The Python remote implementation is not part of the runtime. The Rust API is the single supported implementation surface.
Target implementation contract¶
This page is the issue #536 target contract. Existing lower-level modules remain
public during the migration, but amplihack-cli should call only the
command-shaped API once the api module lands.
The following behaviors are part of the build target:
execandstart_sessionsboth require agent credentials. CLI code readsANTHROPIC_API_KEY, passes it into the API, and validation fails before repository packaging, VM provisioning, or VM pool allocation begins.- Detached sessions use
NODE_OPTIONS=--max-old-space-size=32768. The persisted session state schema includesmemory_mb, and sessions launched throughstart_sessionsrecordmemory_mb: 32768.
High-level command API¶
The api module provides command-shaped request and response types.
use std::path::PathBuf;
use amplihack_remote::{
CommandMode, ExecOptions, ListOptions, OutputOptions, RemoteError,
StartOptions, VMOptions, VMSize,
};
async fn run_remote(repo: PathBuf, api_key: String) -> Result<(), RemoteError> {
let result = amplihack_remote::exec(ExecOptions {
repo_path: repo,
command: CommandMode::Auto,
prompt: "implement remote list JSON output".to_string(),
max_turns: 10,
vm_options: VMOptions::default(),
timeout_minutes: 120,
skip_secret_scan: false,
api_key,
})
.await?;
println!("remote exit code: {}", result.exit_code);
Ok(())
}
CommandMode¶
CommandMode implements string parsing and display using the CLI values
auto, ultrathink, analyze, and fix.
Validation¶
High-level options validate the same rules as the CLI:
| Field | Rule |
|---|---|
prompt |
Must not be empty after trimming whitespace. |
command |
Must be one of auto, ultrathink, analyze, or fix. |
max_turns for exec |
Must be 1..=50. |
timeout_minutes for exec |
Must be 5..=480. |
api_key for exec and start |
Must not be empty. |
start prompts |
At least one non-empty prompt is required. |
output session ID |
Must exist in the state store. |
kill session ID |
Must exist in the state store. |
Validation failures return RemoteError and do not partially update state.
Synchronous execution¶
Use exec for the blocking one-shot workflow.
exec performs:
- Repository, option, and credential validation.
- Context packaging and secret scanning.
- VM provisioning or reuse through azlin.
- Context transfer.
- Remote command execution.
- Log and git-state retrieval.
- Result integration and VM cleanup.
ExecOptions¶
pub struct ExecOptions {
pub repo_path: PathBuf,
pub command: CommandMode,
pub prompt: String,
pub max_turns: u32,
pub vm_options: VMOptions,
pub timeout_minutes: u64,
pub skip_secret_scan: bool,
pub api_key: String,
}
WorkflowResult¶
pub struct WorkflowResult {
pub exit_code: i32,
pub summary: Option<IntegrationSummary>,
pub vm_name: Option<String>,
}
execute_remote_workflow remains available as a compatibility wrapper around
exec for existing Rust callers.
Detached sessions¶
Use detached session APIs for remote start, list, output, kill, and
status.
Start sessions¶
pub struct StartOptions {
pub repo_path: PathBuf,
pub prompts: Vec<String>,
pub command: CommandMode,
pub max_turns: u32,
pub size: VMSize,
pub region: Option<String>,
pub tunnel_port: Option<u16>,
pub api_key: String,
pub state_file: Option<PathBuf>,
}
pub struct StartResult {
pub sessions: Vec<Session>,
}
start_sessions validates all prompts and the API key before side effects begin.
It then packages context separately for each prompt, allocates VM pool capacity,
transfers context, starts a tmux session, and marks each successful session
running.
List sessions¶
Capture output¶
pub struct OutputOptions {
pub session_id: String,
pub lines: u32,
pub state_file: Option<PathBuf>,
}
pub struct SessionOutput {
pub session: Session,
pub output: String,
}
Unknown session IDs return RemoteError::SessionNotFound rather than an empty
string.
Kill sessions¶
pub struct KillOptions {
pub session_id: String,
pub force: bool,
pub state_file: Option<PathBuf>,
}
pub struct KillResult {
pub session: Session,
pub remote_kill_attempted: bool,
pub remote_kill_succeeded: bool,
pub capacity_released: bool,
}
Without force, remote tmux kill failures return an error and leave local state
unchanged. With force, state is updated even when the VM is unreachable.
Status¶
pub struct StatusOptions {
pub state_file: Option<PathBuf>,
}
pub struct RemoteStatus {
pub pool: PoolStatus,
pub sessions: SessionCounts,
pub total_sessions: usize,
}
pub struct SessionCounts {
pub running: usize,
pub completed: usize,
pub failed: usize,
pub killed: usize,
pub pending: usize,
}
VM pooling¶
VMSize maps detached session tiers to Azure SKUs and capacity.
| Variant | Azure SKU | Capacity |
|---|---|---|
S |
Standard_D8s_v3 |
1 |
M |
Standard_E8s_v5 |
2 |
L |
Standard_E16s_v5 |
4 |
XL |
Standard_E32s_v5 |
8 |
let size = VMSize::L;
assert_eq!(size.capacity(), 4);
assert_eq!(size.azure_size(), "Standard_E16s_v5");
Idle VM cleanup¶
VMPoolManager::cleanup_idle_vms reclaims VMs that have no active sessions and
whose grace period has elapsed. It returns the names of the VMs it actually
reclaimed — never the names it merely attempted.
| Parameter | Meaning |
|---|---|
grace_period_minutes |
Minimum idle age before a VM with zero active sessions becomes a cleanup candidate. A VM with no recorded created_at timestamp is always eligible. |
Return value: Vec<String> containing only the names of VMs whose cloud
deallocation was confirmed. A VM that was attempted but not confirmed
deallocated is not included, and it remains in the pool.
Deallocation is verified, never assumed¶
Cleanup delegates to Orchestrator::cleanup(&vm, /* force */ true), which
returns Result<bool, RemoteError>. The cleanup helper matches all three
possible outcomes exhaustively — there is no catch-all fallback that could
silently treat an un-reclaimed VM as removed:
cleanup outcome |
Meaning | Pool action | Included in return Vec |
Log level |
|---|---|---|---|---|
Ok(true) |
Cloud deallocation confirmed | Entry dropped from pool | Yes | info! (batch summary) |
Ok(false) |
cleanup ran but did not confirm deallocation — covers every failure mode (azlin non-zero exit, command-spawn error, or the 120s timeout) |
Entry retained in pool for retry | No | warn! |
Err(e) |
Defensive arm for a hard error returned by cleanup |
Entry retained in pool for retry | No | error! |
force = truecollapses failures intoOk(false).cleanup_idle_vmsalways callscleanup(vm, force = true). Underforce,Orchestrator::cleanupintentionally downgrades every failure (non-zero exit, spawn error, timeout) toOk(false)and never returnsErr. So for this call path the operative outcomes areOk(true)(reclaim) andOk(false)(retain); theErr(e)arm is handled defensively but is not produced whileforceistrue. Do not rely on anOk(false)vsErrsplit to distinguish "ran-but-unconfirmed" from "could-not-run" — both collapse intoOk(false)here.
Whatever the failure mode, a retained VM is picked up again on the next
cleanup_idle_vms pass, so a transient cloud failure self-heals without leaking
a billable resource. This is the safety property tracked by issue #870: a
swallowed cleanup failure previously dropped the VM from tracking while it kept
running (and billing) in Azure.
Logging contract¶
For each retained VM, the cleanup helper emits one structured tracing event
using typed fields (no string interpolation of untrusted output):
WARN cleanup did not confirm deallocation; VM retained for retry vm=<name>
ERROR cleanup failed; VM retained to avoid orphaning billable resource vm=<name> error=<RemoteError Display>
The WARN event is the only one that fires for this call path: because
cleanup_idle_vms calls with force = true, every real failure arrives as
Ok(false) and is logged at warn!. The ERROR event belongs to the defensive
Err(e) arm, which cannot occur while force = true; it exists only to keep the
match exhaustive should the caller ever switch to force = false.
The helper logs only the VM name (warn! case) plus the RemoteError Display
string (error! case), which is a curated cleanup message — never VM tags,
credentials, or command output.
A pass that reclaims at least one VM also emits a single batch summary,
info!(count = <n>, "cleaned up idle VMs"), where <n> is the count of
confirmed reclaims (Ok(true)) only — retained VMs are never counted.
Layered hygiene caveat. This "no raw output" guarantee applies to the pool helper only. The underlying
Orchestrator::cleanupalready logs its ownwarn!("VM cleanup failed: {stderr}")containing the rawazlinstderr on a non-zero exit (orchestrator.rs). That pre-existing log is outside issue870's
vm_poolscope; if rawazlinoutput must be kept out of logs¶entirely, redacting that orchestrator
warn!is a required follow-up. Expect up to two events per unconfirmed cleanup: the orchestrator's stderrwarn!and the pool helper's retentionwarn!.
Operational follow-up¶
Retention failures surface only at warn! and repeat on every subsequent pass
until the VM is finally reclaimed or removed by hand. Since an un-reclaimed VM
keeps billing, operators should alert on repeated retentions for the same vm
rather than eyeballing logs. Emitting a retained-VM gauge/counter from
cleanup_idle_vms (so dashboards and alerts can watch it) is the recommended
follow-up tracked alongside issue #870.
Example¶
// Reclaim VMs idle for 30+ minutes.
let reclaimed = pool.cleanup_idle_vms(30).await;
// `reclaimed` lists only VMs that Azure confirmed as deallocated.
for name in &reclaimed {
println!("reclaimed {name}");
}
// VMs whose cleanup was unconfirmed or failed are still in the pool and
// will be retried on the next call — inspect them via get_pool_status().
let status = pool.get_pool_status();
println!("{} VMs still tracked", status.total_vms);
State persistence¶
State is written (save_state) only when at least one VM was truthfully
reclaimed (!reclaimed.is_empty()). When every candidate is retained, the pool
is unchanged and no write occurs.
State management¶
SessionManager owns session lifecycle state.
let mut manager = SessionManager::new(None)?;
let session = manager.create_session(
"amplihack-azureuser-20260502",
"implement remote status JSON",
Some("auto"),
Some(10),
Some(32_768),
)?;
manager.start_session(&session.session_id)?;
Detached sessions created by start_sessions store Some(32_768) for
memory_mb. This value is part of the persisted state contract and mirrors the
remote tmux environment's NODE_OPTIONS=--max-old-space-size=32768 setting.
Session state values are serialized as lowercase strings:
All state writes use the state_lock module so concurrent processes update the
same state file safely.
Errors¶
Public APIs return RemoteError.
pub enum RemoteError {
Packaging(ErrorContext),
Provisioning(ErrorContext),
Transfer(ErrorContext),
Execution(ErrorContext),
Integration(ErrorContext),
Cleanup(ErrorContext),
SessionNotFound { session_id: String },
Validation(ErrorContext),
}
Each error carries a user-facing message and optional context. CLI handlers map errors to process exit codes:
| Error | Exit code |
|---|---|
SessionNotFound |
3 |
| Validation, packaging, provisioning, transfer, execution, integration, cleanup | 1 |
| CLI parser errors before API dispatch | 2 |
Testing seams¶
Remote infrastructure calls are behind traits so unit tests do not require live Azure resources:
#[async_trait::async_trait]
pub trait RemoteCommandRunner {
async fn azlin(&self, args: &[String]) -> Result<CommandOutput, RemoteError>;
async fn connect(&self, vm_name: &str, command: &str) -> Result<CommandOutput, RemoteError>;
async fn transfer(&self, source: &Path, vm_name: &str, destination: &Path)
-> Result<(), RemoteError>;
}
The production runner shells out to azlin. Tests use an in-memory runner to
assert command construction, state transitions, output parsing, force-kill
behavior, and JSON response shape.