Skip to content

azlin disk check

Report whether a VM's data disks are provisioned the way azlin intended. Read only — it never formats, mounts, or writes anything on the VM.

Usage

azlin disk check VM_NAME [OPTIONS]

Arguments

  • VM_NAME - VM name or session name (required)

Options

Option Description
--resource-group, --rg TEXT Azure resource group (default: configured resource group)
--json Emit machine-readable JSON instead of the table
-h, --help Show help message

Exit codes

azlin disk check is meant to be usable in a script or a cron job, so the verdict is in the exit status, not only in the text:

Code Meaning
0 Every expected disk is healthy, or the VM has no azlin data disks
1 Degraded — at least one disk is not at the healthy stage
2 The check could not be completed (VM unreachable, probe output unparseable)

Code 2 is never reported as healthy. An unreachable VM is an unknown VM, and this command will not answer a question it could not ask.

1 is also azlin's generic failure status, and clap exits 2 on a usage error, so the codes alone cannot tell "this VM is degraded" apart from "the command did not run". A script that needs the difference should read --json and branch on status, which is only ever emitted by a check that completed.

Examples

A correctly provisioned VM

azlin disk check build-vm

Output:

VM: build-vm  (rg: azlin-rg)
Storage: ok

ROLE  LUN  DEVICE    SIZE  STAGE
  home  0    /dev/sdb  100G  healthy
  tmp   1    /dev/sdc  64G   healthy

Provisioning: complete, status=ok

Exit code 0.

A VM whose disks were never set up

This is the shape of issue #1131: both disks attached, neither formatted, everything running on the 30 GB OS disk.

azlin disk check dev

Output:

VM: dev  (rg: rysweet-linux-vm-pool)
Storage: degraded

  ROLE  LUN  DEVICE    SIZE   STAGE
  home  0    /dev/sdb  1000G  raw
        no filesystem on the device; /home/<user> is on the OS disk
  tmp   1    /dev/sdc  200G   raw
        no filesystem on the device; /tmp is on the OS disk

Provisioning: complete, status unknown (no ledger — VM predates it)

Repair in place with:  azlin disk repair dev

Exit code 1.

The stage is decided from what is on the VM right now — the LUN symlinks, blkid, and the mount table — not from the provisioning ledger. VMs created before the ledger existed have no ledger, and they are exactly the VMs most likely to be broken. When a ledger is present it is reported as corroborating detail, and the failed section names are listed:

Provisioning: complete, status=degraded
  failed sections: apt-update, apt-install

Checking every VM in a resource group

for vm in $(azlin -o csv list | tail -n +2 | cut -d, -f1); do
  azlin disk check "$vm" >/dev/null 2>&1 || echo "$vm: storage degraded"
done

Output:

dev: storage degraded
deva2: storage degraded
deva3: storage degraded

JSON output

azlin disk check dev --json

Output:

{
  "vm": "dev",
  "resource_group": "rysweet-linux-vm-pool",
  "status": "degraded",
  "disks": [
    {
      "role": "home",
      "lun": 0,
      "device": "/dev/sdb",
      "size_gb": 1000,
      "stage": "raw",
      "detail": "no filesystem on the device; /home/<user> is on the OS disk"
    },
    {
      "role": "tmp",
      "lun": 1,
      "device": "/dev/sdc",
      "size_gb": 200,
      "stage": "raw",
      "detail": "no filesystem on the device; /tmp is on the OS disk"
    }
  ],
  "provisioning": {
    "complete": true,
    "status": "unknown",
    "ledger_present": false,
    "failed_sections": []
  }
}

status is one of ok, degraded, no-disks, or unknown. stage is one of absent, raw, formatted, backing-mounted, healthy — see Provisioning stages.

detail says /home/<user> literally rather than naming the account. The probe output does not carry the admin username, and the parser that writes this field does not invent one.

How the check works

azlin disk check opens one SSH session (over Bastion if that is how the VM is reached) and runs a read-only probe that prints facts, one line per expected disk:

azlin-disk lun=0 role=home dev=/dev/sdb size=107374182400 fstype=ext4 label=azlin-home backing=yes bind=yes
azlin-disk lun=1 role=tmp dev=/dev/sdc size=68719476736 fstype= label= backing=no bind=no
azlin-provisioning complete=yes status=ok ledger=yes failed=

One azlin-disk line per expected disk, then exactly one azlin-provisioning line:

  • dev comes from readlink -f /dev/disk/azure/scsi1/lunN, never from /dev/sd* guessing. The LUN symlink is the stable identity — /dev/sdb can name a different disk after a reboot — which is why lun and dev are reported as separate fields and why the DEVICE column is only ever a resolved kernel name. To address the disk yourself, use the lun.
  • size is the raw byte count from lsblk -bdno SIZE on that resolved device, rendered as the SIZE column and as size_gb in JSON. It is read from the device rather than from the Azure disk record deliberately: a disk that is attached in Azure but has no device on the VM is exactly the absent case, and the two sources would disagree there. A disk at stage absent has no device, so it has no size — -- in the table, null in JSON.
  • fstype and label come from blkid; empty means the disk is raw
  • backing and bind come from findmnt, falling back to /proc/mounts
  • when a target appears more than once in the mount table, the last exact target match is authoritative. Linux stacks a later mount over an earlier one, so this is the mount that receives a write and the one df reports. The same rule applies whether the probe reads findmnt or /proc/mounts. Earlier records remain visible in the mount table but do not determine the disk stage.
  • the azlin-provisioning line reads /var/lib/azlin/provisioning-complete, /var/lib/azlin/provisioning-status, and the section names of any failed rows in /var/lib/azlin/provisioning.tsv. On a VM that predates the ledger the files are absent, so the probe emits ledger=no status=unknown — a first-class case, not a parse failure, and the common one for the fleet this command exists to fix.

azlin decides the verdict; the probe has no opinion. Output it cannot parse — an older image, a truncated session — yields unknown and exit code 2, never a false degraded. A missing azlin-provisioning line is a parse failure; a line reporting ledger=no is not.

Stacked mounts use the effective record

A correctly provisioned /tmp can have an older tmpfs mount below the data-disk bind. It can also contain duplicate data-disk records left by an earlier mount operation:

tmpfs    on /tmp             type tmpfs (rw,nosuid,nodev,...)
/dev/sdb on /home/azureuser  type ext4 (rw,relatime)
/dev/sdc on /mnt/tmp-data    type ext4 (rw,relatime)
/dev/sdc on /tmp             type ext4 (rw,relatime)
/dev/sdc on /tmp             type ext4 (rw,relatime)

For /tmp, the final /dev/sdc record is the effective mount. The tmp role is therefore healthy, not backing-mounted, provided /dev/sdc is the device resolved from the role's LUN. The first tmpfs record is not evidence that /tmp is still on the OS disk.

The source comparison accepts the exact resolved device and the bind-source form emitted by findmnt, such as /dev/sdc[/tmp]. A bracketed suffix is valid only when it begins immediately after the exact device, contains a non-empty absolute path beginning with /, and ends with the closing ] at the end of the source. Similar-looking devices such as /dev/sdc1, incomplete suffixes such as /dev/sdc[/tmp, empty suffixes such as /dev/sdc[], and trailing content such as /dev/sdc[/tmp]extra do not match.

Several records for the same exact target are not ambiguous: the probe selects the last one and evaluates only that effective record. unknown is reserved for output that cannot identify and parse a complete source-target record, not for an ordered stack of duplicate exact-target records.

The probe is cheap and read-only, which is why the same result also feeds the Storage column of azlin list --with-health and of azlin health.

Storage in the health surfaces

Two surfaces report the same verdict without you having to ask per VM, both from this probe and both over the same connection they use for their other metrics — so the storage question costs a round trip, not a second sweep.

azlin list --with-health

Output:

SESSION   OS       Status    IP            Region    CPU  Mem   Agent  CPU%  Mem%  Disk%  Storage
dev       Ubuntu   running   10.0.1.4      westus2   8    32G   ok     12    41    98     degraded
build-vm  Ubuntu   running   10.0.1.7      westus2   8    32G   ok     4     22    31     ok
scratch   Ubuntu   running   10.0.1.9      westus2   4    16G   --     --    --    --     --

azlin health

Output:

Health Dashboard — Four Golden Signals (azlin-rg)
┌────────────────────┬──────────┬──────────┬──────┬──────┬────────┬──────┬─────────┐
│ VM Name            │ State    │ Agent    │Errors│ CPU %│Memory %│Disk %│ Storage │
├────────────────────┼──────────┼──────────┼──────┼──────┼────────┼──────┼─────────┤
│ dev                │ Running  │ OK       │     0│  12.0│    41.0│  98.0│ degraded│
│ build-vm           │ Running  │ OK       │     0│   4.0│    22.0│  31.0│ ok      │
│ scratch            │ Running  │ --       │    --│    --│      --│    --│ --      │
└────────────────────┴──────────┴──────────┴──────┴──────┴────────┴──────┴─────────┘

-- means the probe did not run or could not be parsed — the same convention the other health columns use. It is not a pass. A VM with no azlin data disks also shows --: there is no storage layout for it to be ok about.