# Permadyn AI: sample hardware acceptance plan Blank planning template. This is not a benchmark result, customer case study, or certification. Reviewed: September 15, 2026. ## 1. Define the decision - Application and representative tasks: - Business owner and technical reviewer: - Equipment budget and separate implementation/support budget: - Proposed hardware and alternative to compare: - Procurement decision date and supplier quote expiry: - Acceptance thresholds agreed before testing: ## 2. Pin the complete configuration - Hardware model, number of systems, memory per device, storage: - Network topology, switch, cables, and measured link behavior: - Power, cooling, physical location, and recovery access: - Operating system, driver, inference runtime, and container versions: - Model repository, exact revision, license, and quantization: - Context length, cache precision, output length, and batching: - Per-node model placement and memory reserve: - Retrieval, embeddings, reranking, and application processes: Record each device separately. Aggregate installed memory is not a single usable memory pool. ## 3. Evaluate quality - Approved, non-sensitive evaluation set and version: - Task-specific scoring rubric and expected outputs: - Baseline system and minimum acceptable result: - Reviewers, disagreement handling, and failure examples: - Comparison of quantized and higher-precision behavior where relevant: - Unsupported questions, source citations, and escalation behavior: Report the tested sample size, scoring method, and limitations. Do not substitute vendor benchmarks for application results. ## 4. Measure application performance For each concurrency and context setting, record repeated runs: | Configuration | Context | Concurrent requests | Time to first token p50/p95 | End-to-end latency p50/p95 | Output tokens/sec | Peak memory per device | Error rate | | --- | --- | --- | --- | --- | --- | --- | --- | | To be tested | | | | | | | | - Separate cold starts from warm runs. - State whether prefix caching is enabled and whether prompts share a prefix. - Report input-processing speed separately from output generation. - Record test duration, request count, queue delays, and measurement method. - For document/image/speech workloads, add the appropriate task throughput and quality metrics. - Record wall-power measurements only when measured with an identified method. ## 5. Verify the complete system - Authentication, authorization, and source-access restrictions: - Permitted outbound connections and verified local-data boundary: - Tool-action validation and human approval where required: - Unavailable integrations, interrupted requests, and retry behavior: - Model/runtime update and rollback: - Node/link failure behavior for multi-node deployments: - Backup restore and service recovery rehearsal: - Monitoring, alert routing, incident ownership, and support coverage: A cluster is not automatically highly available. Test the actual application behavior when a node or link fails. ## 6. Record the decision and handoff - Measured results against each agreed threshold: - Remaining issues and accepted limitations: - Recommendation: buy / change configuration / retest / defer: - Final bill of materials, supplier quote, and lead time: - Source-code rights, accounts, and infrastructure ownership: - Operating documentation and named support responsibilities: - Reviewer approval and date: Published by Permadyn AI. https://permadyn.ai/services/on-premises-ai