Skip to content

Model-neutrality documentation review evidence

Evidence status and acceptance specification

The reference and usage guide describe the finished-state acceptance contract. Recursive fail-explicit loading and metadata support were pending at the original review and are now implemented in this branch. Supporting-YAML and isolated full-corpus publication now have dedicated regressions; broader asset checks remain implementation specifications; the evidence below identifies only checks actually run. Historical results below retain their original scope and do not validate a subsequently reconciled revision.

The parent skill-load-tests.log records 11 passing name tests and four passing frontmatter type tests. The parent real-skill-loader-tests.log records a failing real-bundle production-loader regression: nine identities were omitted, including all seven nested skills, amplihack-expert, and dynamic-debugger. The latter two skills carry structured token_budget mappings. These observations establish why YAML syntax success is insufficient; they do not establish a corrected loader result.

Final implementation evidence records exact 130-name and 130-path equality, nonempty bodies, directory resolution, metadata and structure results, supporting-YAML counts and exclusions, isolated publication results, neutrality and mirror checks, review outcomes, and tested revision. SDK compilation and runtime checks remain separate. Commit, push, PR URL, and PR check results are recorded only after those actions occur.

This report records the documentation refinement step. Current behavior and the intended configuration contract are described in the reference and the usage guide; configuration validation is checked statically in the target worktree.

Existing work preserved

The 71-file statement described an earlier staged snapshot; it does not describe the later 76-path reconciliation manifest. Those are different snapshots, not interchangeable totals. The copied root snapshot contains 76 tracked changed paths (listed in /d0/ryan/sunfixamp-model-neutrality/refinement-source-manifest.txt); adding the preserved untracked regression and loader fix gives 78 paths in this branch. The current refinement copied the root tracked diff into the clean feat/issue-1533-user-explicitly worktree and preserved the untracked real-bundle regression. Root edits remain untouched. The three refined documentation pages were copied from feat/issue-1532-complete-model.

Scope and decisions

Filesystem enumeration confirms 130 canonical skills (123 top-level, seven nested), 537 accessible canonical text paths, and 323 mirror text paths containing 81 mirrored skills. All paths were read to establish text coverage; this step does not claim a fresh semantic audit of all their contents. The inventory below identifies each canonical skill and mirror availability.

The documentation specifies removal of execution selections, runtime inheritance where supported, explicit required provider configuration, distinct primary/fallback inputs, Azure deployment semantics, and the intended pre-client validation of absent, empty, whitespace-only, and unresolved placeholder values in all three C# examples. The examples and their mirrors reject absent, empty, whitespace-only, and angle-bracket placeholder values with string.IsNullOrWhiteSpace and delimiter checks before constructing clients. API parameter names and credentials remain provider-specific. Historical research, poetic forms, and reviewed API context remain allowed; exact exceptions require reasons and cannot cover whole files.

Directory links from docx and pptx point to canonical common/ooxml; outside-in-testing/examples points to qa-team/examples. Directory links are not recursively followed. The accessible outside-in-testing/README.md file link is read and counted. Its scripts and tests links point to missing qa-team directories; no contents behind those links were available for review.

Validation evidence and limits

Earlier statements about whitespace, links, fences, mirror checks, and caller-directory independence lacked revision, exit-code, and log attachments. They remain historical reports, not reproducible evidence for this branch. Parent frontmatter logs likewise omit a recorded tested revision; they are baseline evidence only. Baseline loader failure is retained at /d0/ryan/sunfixamp-model-neutrality/real-skill-loader-tests.log (nine omitted identities).

Current production-loader validation checks exactly 130 recursive files, exact name equality, resolved source-directory equality, and nonempty bodies. Unit tests cover optional description, extension round trips, typed metadata rejection, duplicate names, malformed delimiters, and empty bodies. Extension preservation does not add execution semantics.

At this historical revision, supporting-YAML, corpus structure/assets, and complete isolated installation coverage were still planned. The follow-up evidence below supersedes the YAML and installation limitations. SDK compilation and provider execution remain unverified; C# configuration rejection is already implemented in all three examples and mirrors, so it is not pending implementation.

Builds use the authorized existing cache /d0/ryan/amplihack-builds/target/amp; temporary files use /d0/ryan/sunfixamp-model-neutrality/tmp. Evidence for this refinement is recorded below with exact commands and log paths. The PR remains open and unmerged.

Recorded refinement validation

Tested tree: base revision 4ccd1977a54e06a7f940507ebc2041f2f638869e plus this branch’s uncommitted 78-path change set, subsequently committed for publication. Commands ran from feat/issue-1533-user-explicitly, with CARGO_TARGET_DIR=/d0/ryan/amplihack-builds/target/amp and TMPDIR=/d0/ryan/sunfixamp-model-neutrality/tmp.

Command Exit Log
cargo test -p amplihack-domain-agents --test bundled_skill_catalog 0 /d0/ryan/sunfixamp-model-neutrality/loader-final.log
cargo test -p amplihack-domain-agents --lib skill_catalog 0 /d0/ryan/sunfixamp-model-neutrality/metadata-final.log
cargo test -p amplihack --test skill_frontmatter_name --test skill_frontmatter_type --test issue_849_skill_mirror_citation 0 /d0/ryan/sunfixamp-model-neutrality/frontmatter-final.log
cargo test -p amplihack-cli issue_1438_skill_publication 0 /d0/ryan/sunfixamp-model-neutrality/publication-final.log
bash amplifier-bundle/recipes/tests/test-skill-model-neutrality.sh 0 /d0/ryan/sunfixamp-model-neutrality/neutrality-final.log
PYTHONDONTWRITEBYTECODE=1 python3 amplifier-bundle/recipes/tests/test_skill_model_neutrality.py 0 /d0/ryan/sunfixamp-model-neutrality/regressions-final.log
bash amplifier-bundle/recipes/tests/test-issue-962-skill-mirror-parity.sh 0 /d0/ryan/sunfixamp-model-neutrality/mirror-final.log
bash -n amplifier-bundle/recipes/tests/test-skill-model-neutrality.sh 0 /d0/ryan/sunfixamp-model-neutrality/syntax-final.log
git diff HEAD --check 0 /d0/ryan/sunfixamp-model-neutrality/whitespace-final.log
shellcheck amplifier-bundle/recipes/tests/test-skill-model-neutrality.sh 0 /d0/ryan/sunfixamp-model-neutrality/shellcheck-final.log

The loader regression passed for all 130 names and resolved paths; 11 catalog unit tests passed, 11 frontmatter name tests, four type tests, six citation tests, one existing installation test, and six neutrality regression tests passed. The neutrality guard reports 130 skills and 860 text paths. Documentation local-link and paired-fence checks exited 0 (/d0/ryan/sunfixamp-model-neutrality/documentation-final.log); the script reads only the three refined pages. Root preservation was checked by cmp of tmp/root.patch and tmp/root-after.patch (exit 0).

Review and double-check examined loader discovery, error propagation, metadata serialization, typed operational fields, exact corpus coverage, configuration validation and documentation claims. The citation target names were confirmed against Cargo registrations; no command rename was warranted. Remaining planned tests are disclosed above. Changing SkillMeta.token_budget from Option<u32> to Option<TokenBudget> is a Rust API change: callers matching scalar budgets must use TokenBudget::Scalar. Repository callers compile in the completed tests.

Pre-commit initially exited 1 for a Rust formatting mismatch (precommit-final.log); cargo fmt -p amplihack-domain-agents corrected it. cargo fmt --all --check then exited 0 (format-final.log), and Clippy passed. The staged-tree pre-commit recheck exited 0 and is logged at /d0/ryan/sunfixamp-model-neutrality/precommit-recheck.log.

CI on commit 740b133b83ce3aa18e844fc227acb54f478d86a4 rejected the copied .py asset under the repository’s tracked-Python guard. The six regression cases were moved into test-skill-model-neutrality-regressions.sh, initially using the existing guard’s inline-Python shell convention. The shell regression command passed all six tests (exit 0, /d0/ryan/sunfixamp-model-neutrality/regressions-shell.log). scripts/check-no-python-assets.sh, ShellCheck on the regression shell, and staged pre-commit all exited 0 (python-assets-final.log, shell-regressions-check.log, precommit-shell.log in the same evidence directory). The guide now names that entry point; the earlier Python command above is historical evidence.

Canonical inventory

“Enumerated” records filesystem coverage, not a semantic neutrality pass.

Canonical skill path Coverage Mirror
agent-generator-tutor Enumerated No
agentic-workflow-first Enumerated No
amplihack-expert Enumerated No
anthropologist-analyst Enumerated Yes
aspire Enumerated No
authenticated-web-scraper Enumerated No
auto-drive-to-merge Enumerated Yes
awesome-copilot-sync Enumerated No
azure-admin Enumerated Yes
azure-devops Enumerated No
backlog-curator Enumerated Yes
biologist-analyst Enumerated Yes
cascade-workflow Enumerated Yes
chemist-analyst Enumerated Yes
claude-agent-sdk Enumerated Yes
code-atlas Enumerated No
code-philosophy Enumerated No
code-smell-detector Enumerated Yes
code-visualizer Enumerated Yes
collaboration/creating-pull-requests Enumerated Yes
computer-scientist-analyst Enumerated Yes
consensus-voting Enumerated Yes
context-management Enumerated No
crusty-old-engineer Enumerated Yes
cybersecurity-analyst Enumerated Yes
debate-workflow Enumerated Yes
default-workflow Enumerated Yes
dependency-resolver Enumerated No
design-patterns-expert Enumerated Yes
dev-orchestrator Enumerated No
development/architecting-solutions Enumerated Yes
development/setting-up-projects Enumerated Yes
documentation-writing Enumerated Yes
docx Enumerated Yes
dotnet-exception-handling Enumerated No
dotnet-install Enumerated No
dotnet10-pack-tool Enumerated No
dynamic-debugger Enumerated Yes
e2e-outside-in-test-generator Enumerated No
economist-analyst Enumerated Yes
email-drafter Enumerated Yes
engineer-analyst Enumerated Yes
environmentalist-analyst Enumerated Yes
epidemiologist-analyst Enumerated Yes
ethicist-analyst Enumerated Yes
eval-recipes-runner Enumerated Yes
fleet Enumerated No
fleet-copilot Enumerated No
futurist-analyst Enumerated Yes
gh-aw-adoption Enumerated No
gh-work-report Enumerated No
gherkin-expert Enumerated Yes
github Enumerated No
github-copilot-cli Enumerated No
github-copilot-cli-expert Enumerated No
github-copilot-sdk Enumerated No
goal-seeking-agent-pattern Enumerated Yes
historian-analyst Enumerated Yes
indigenous-leader-analyst Enumerated Yes
investigation-workflow Enumerated Yes
journalist-analyst Enumerated Yes
knowledge-extractor Enumerated Yes
lawyer-analyst Enumerated Yes
learning-path-builder Enumerated Yes
lsp-setup Enumerated No
markitdown Enumerated No
mcp-manager Enumerated Yes
meeting-synthesizer Enumerated Yes
merge-ready Enumerated No
mermaid-diagram-generator Enumerated Yes
meta-cognitive/analyzing-deeply Enumerated Yes
microsoft-agent-framework Enumerated Yes
migrate Enumerated No
model-evaluation-benchmark Enumerated Yes
module-spec-generator Enumerated Yes
multi-repo Enumerated No
multitask Enumerated No
n-version-workflow Enumerated Yes
novelist-analyst Enumerated Yes
npe-hunting-workflow Enumerated No
outside-in-testing Enumerated Yes
oxidizer-workflow Enumerated No
pdf Enumerated Yes
philosopher-analyst Enumerated Yes
philosophy-compliance-workflow Enumerated Yes
physicist-analyst Enumerated Yes
pm-architect Enumerated Yes
poet-analyst Enumerated Yes
political-scientist-analyst Enumerated Yes
pptx Enumerated Yes
pr-guide Enumerated No
pr-review-assistant Enumerated Yes
pre-commit-manager Enumerated No
property-based-testing Enumerated No
psychologist-analyst Enumerated Yes
qa-team Enumerated Yes
quality/reviewing-code Enumerated Yes
quality/testing-code Enumerated Yes
quality-audit Enumerated Yes
remote-work Enumerated Yes
repository-oom-audit Enumerated No
research/researching-topics Enumerated Yes
roadmap-strategist Enumerated Yes
self-improving-agent-builder Enumerated No
session-learning Enumerated No
session-replay Enumerated No
session-to-agent Enumerated No
shadow-testing Enumerated No
signal Enumerated No
signal-setup Enumerated No
silent-degradation-audit Enumerated No
skill-builder Enumerated Yes
smart-test Enumerated No
sociologist-analyst Enumerated Yes
socratic-review Enumerated No
statler-waldorf Enumerated No
storytelling-synthesizer Enumerated Yes
supply-chain-audit Enumerated No
test-gap-analyzer Enumerated Yes
tla-plus-expert Enumerated Yes
transcript-viewer Enumerated Yes
ultrathink-orchestrator Enumerated Yes
urban-planner-analyst Enumerated Yes
verus-expert Enumerated Yes
work-delegator Enumerated Yes
work-iq Enumerated No
workflow-enforcement Enumerated No
workiq-wsl Enumerated No
workstream-coordinator Enumerated Yes
xlsx Enumerated Yes

The follow-up replaces the regression heredoc with native Bash fixtures rather than embedding the former Python test asset. The same six integration scenarios remain, including seven provider recommendation inputs; fixtures are isolated under TMPDIR. The production audit still uses its existing inline Python detector. Both no-Python gates and ShellCheck pass. Supporting YAML now has a Rust corpus test which parses every YAML document without executing example commands.

Follow-up local results: production catalog acceptance passes for 130 identities, canonical paths and nonempty bodies, including seven nested skills. Rust parses all 24 supporting YAML files and every document, with zero exclusions; a malformed second-document regression also passes. Isolated real-corpus publication verifies all 130 published SKILL.md files byte-for-byte. The issue #1438 local installation fixture and issue #1277 nested support-file fixture pass. Frontmatter name/type checks pass 15 tests, citation checks pass six, and issue #962 passes six checks. Logs use the native- prefix under /d0/ryan/sunfixamp-model-neutrality; native-review.md records review and double-check findings. Broad validation of all referenced assets and provider SDK execution remain unverified.