On-premises AI deployments

Use AI without sending your business data to an outside AI service. We build assistants, document tools, and workflows that run on your own hardware. We connect them to your systems, test them before release, and support them as agreed.

ScopeFrom one local application to a shared internal AI platform
Delivery timingAgreed after a workload, hardware, access, and security review
How we check itTask quality, access controls, outbound traffic, load, and recovery, tested before you accept the system

AI applications on your own hardware

AI means the models, and the application built around them, run on hardware at your site. We select local large language models () and task-specific models, connect the data sources you approve, and build the application. Local hardware alone does not prove that data stays inside the agreed boundary, so we test the complete system before acceptance.

Good fit

  • Organizations whose contracts or policies prohibit processing data outside their own systems
  • Teams handling confidential records, intellectual property, or internal research
  • Businesses that need AI in offline or tightly restricted environments

Not the right fit

  • Tasks that ordinary software, a rule, or a process change can handle
  • Projects that require a model offered only as a hosted service, with no acceptable local alternative
  • Workloads where an approved hosted service meets the requirements at a lower total cost of ownership
  • Deployments with no one responsible for the infrastructure or the application

Custom hardware solutions

What can you run on your own hardware?

The hardware you need depends mainly on three things: the model, how much text it handles in one request (its ), and how many people use it at once. The tiers below are reference configurations, not fixed packages or stock listings. Custom-configured means we size the system to your workload and confirm , memory, storage, and model compatibility before purchase.

Based on current hardware pricing as of , subject to change. Prices are planning budgets in US dollars, not supplier quotes. Configuration and availability affect the final cost. Tax, delivery, deployment, and ongoing application support are separate.

Subject to availability. GPUs and complete systems may be backordered. Supplier lead times, configuration, testing, and shipping all affect when equipment arrives. We confirm availability and expected dates before ordering. If a part must be substituted, we recheck compatibility and ask for your approval first.

In this guide: three budget picks, DGX Spark clusters, model specifications and hardware fit, and a comparison with hosted AI.

Three budget picks

Our default picks for text-first , which means using an already trained model to answer mostly text requests. The right choice is the smallest system that passes quality and load tests on your own tasks.

$2,000–$5,000Equipment budget bracket

Focused pilot

NVIDIA RTX 5060 Ti

Our pick for one narrow text workflow: a workstation with the 16 version of this card.

Model to test first: Qwen3 8B at 8K .

Move up a tier when the chosen model, its context, or the number of simultaneous requests no longer fits.

See pros and cons

$5,000–$15,000Equipment budget bracket

Versatile application build

NVIDIA RTX 5090

Our general-purpose pick: 32 GB of in a single- deployment. Plan $7,000–$12,000 for the configured workstation.

Model to test first: Qwen3 30B-A3B at 16K context.

Choose DGX Spark instead when fitting a larger model matters more than speed. It holds more, but its lower memory limits how fast it generates text.

See pros and cons

$15,000–$35,000Equipment budget bracket

Larger-model deployment

NVIDIA RTX PRO 6000 Blackwell

Our pick for larger models on one GPU, with 96 GB of GPU memory. Plan $20,000–$35,000 for the configured workstation.

Model to test first: Llama 3.3 70B Instruct at 32K context.

Consider a cluster only when the model’s size or separate services make it worth running across several machines.

See pros and cons

These picks and the pros and cons below are our reading of published specifications and model requirements, not benchmark results.

Entry-level workstation

NVIDIA RTX 5060 Ti

$2,000–$5,000Planning budget
16

Best fit: A focused pilot, an extraction pipeline, or a small local assistant.

Pros

  • The lowest budget in this guide, and a practical way to test one narrow use case.
  • 16 GB of GPU memory holds the 4B and 8B examples with room for their working context.

Cons

  • Limited room for larger models, long conversations, or simultaneous requests.
  • Order the 16 GB card specifically. The 8 GB version is a different configuration, and these sizing examples do not apply to it.

RTX 5060 Ti specifications.

AI workstation

NVIDIA RTX 5090

$7,000–$12,000Planning budget
32 GB GPU memory

Best fit: A versatile CUDA workstation for applications whose models fit within 32 GB.

Pros

  • High memory bandwidth on a single GPU, which keeps the deployment simpler than a distributed system.
  • Holds the small, medium, and vision-model examples without splitting weights across machines.

Cons

  • 32 GB is a hard limit: the 70B example needs more memory.
  • The RTX 5090 is a high-power card. The full workstation needs a specified power supply, cooling, physical space, and service coverage.

RTX 5090 specifications.

DGX Spark desktop

NVIDIA DGX Spark (GB10)

$4,700–$7,000Planning budget
128 GB

Best fit: Testing larger models where memory capacity and a compact footprint matter more than top speed for a single request.

Pros

  • 128 GB of unified memory in a compact CUDA-capable system.
  • Built-in ConnectX-7 networking supports NVIDIA’s documented multi-Spark configurations.

Cons

  • Memory bandwidth is 273 GB/s, much lower than the discrete GPUs listed here. More memory does not mean faster text generation.
  • The Arm CPU needs compatible containers and dependencies, and memory is shared with the operating system and application.

DGX Spark specifications; NVIDIA’s Spark price update.

Apple silicon desktop

Apple M5 Ultra

$6,000–$12,000Planning budget
96 GB unified memory

Best fit: A desktop deployment built on inference software that supports Apple silicon.

Pros

  • A large unified-memory configuration can keep the larger example models on one machine.
  • Metal and MLX provide ways to run models locally on Apple hardware.

Cons

  • CUDA-only software must be replaced or ported; a working NVIDIA deployment does not carry over unchanged.
  • Memory is chosen at purchase, and the operating system and other applications share the same pool.

Mac Studio specifications; Apple’s pricing and availability announcement.

Professional AI workstation

NVIDIA RTX PRO 6000 Blackwell

$20,000–$35,000Planning budget
96 GB GPU memory

Best fit: A larger-model application that benefits from one high-memory GPU and professional workstation support.

Pros

  • 96 GB of GPU memory with ECC and high memory bandwidth.
  • Keeps many larger models on one GPU instead of splitting them across machines.

Cons

  • A large hardware investment: buy it after workload testing, not before.
  • Power, cooling, chassis compatibility, and contracted system support still need to be specified.

RTX PRO 6000 specifications.

Multi-GPU system

2 × NVIDIA RTX PRO 6000 Blackwell

$40,000–$65,000+Planning budget
192 GB total GPU memory

Best fit: A model too large for one GPU, or separate model services sharing one larger machine.

Pros

  • 192 GB of installed GPU memory across two cards.
  • Can split a compatible model across both GPUs or give each card its own service.

Cons

  • The two cards’ memory is not automatically pooled; model placement, runtime support, and PCIe topology all matter.
  • More power, cooling, and operating complexity. Two GPUs do not guarantee twice the throughput.

RTX PRO 6000 specifications.

From custom hardware to a working deployment

We specify the configuration, test models on your tasks, connect your data and software, and set up permissions and monitoring before handing the system over. Deployment dates depend on equipment delivery and site readiness.

Acceptance plan

Agree on the test before you buy hardware

Our proposal names what will be tested, who reviews the results, and what result qualifies the system for purchase and release.

Download a sample acceptance planBlank evaluation template, not a completed benchmark report.
  1. Record the exact configurationHardware, model revision, , , context, and deployment topology.
  2. Measure your workloadTask quality, time to first response, sustained , concurrent requests, and peak memory.
  3. Verify the complete deploymentPermissions, integration failures, recovery, and the people responsible for operating the release.

DGX Spark clusters

When one DGX Spark is not enough

Several DGX Spark systems, or nodes, can be connected. Compatible software can split one large model across the nodes, or each node can run its own models and services. We compare clusters of 2, 3, 4, and 8 nodes below; the right layout depends on your workload.

Systems-only figures use NVIDIA’s published price of $4,699 per DGX Spark Founders Edition, reviewed September 15, 2026. Switches, cables, installation, tax, delivery, and support are extra.

Installed memory totals add up separate machines; they are not one shared pool. Each node needs room for its share of the model , plus , the operating system, and overhead. Adding nodes does not guarantee a proportional speed increase, or that another node takes over automatically when one fails.

Documented topology

Two-Spark pair

Direct QSFP connection

Systems only
$9,398
Installed
256 total

Best fit: A compatible model too large for one Spark, or two independent AI services.

Pros

  • Adds model capacity without a high-speed switch.
  • A documented first step into multi-node deployment.

Cons

  • A split model depends on both nodes and the link between them.
  • Communication overhead can cancel out the extra computing power.
NVIDIA two-Spark playbook

Documented topology

Three-Spark cluster

Direct cabling, without a switch

Systems only
$14,097
Installed unified memory
384 GB total

Best fit: Testing a larger sharded model, or separating inference, evaluation, and batch work.

Pros

  • More capacity within NVIDIA Sync’s direct-connect limit.
  • More options for placing independent services.

Cons

  • Requires the documented cabling and runtime configuration.
  • The model may not divide evenly across three nodes.
NVIDIA three-Spark playbook

Documented topology

Four-Spark cluster

Compatible high-speed Ethernet switch

Systems only
$18,796
Installed unified memory
512 GB total

Best fit: A compact multi-node deployment with a defined serving or evaluation workload.

Pros

  • Within NVIDIA Sync’s documented switched-cluster limit.
  • Can spread a compatible model across nodes or run separate copies.

Cons

  • The switch, cables, power, and cooling add to the node cost.
  • Throughput and recovery still need end-to-end testing.
NVIDIA switched-cluster playbook

Custom engineering review

Eight-Spark cluster

Custom switched topology

Systems only
$37,592
Installed unified memory
1,024 GB total

Best fit: A feasibility study for independent serving groups, batch workloads, or a validated distributed runtime.

Pros

  • More nodes can host separate model copies or workload groups.
  • Room to explore beyond a single four-node deployment.

Cons

  • Beyond NVIDIA Sync’s documented four-node setup limit; not a turnkey supported configuration.
  • We validate networking, orchestration, and model support before ordering. No eight-node performance is claimed.
NVIDIA clustering and support limits
Enough memory is not a proven fit

The version of Qwen3 235B-A22B is larger than one Spark’s memory. A pair has enough combined memory to try splitting it across both machines, but that does not show that it will fit or how fast it will run. For larger clusters, we test the actual model split, network, simultaneous load, and failure behavior before you commit to the equipment.

Plan a DGX Spark cluster

model reference

Compare model size, context, and licensing

Open-weight models can be downloaded and run on your own hardware. The models below are representative examples with downloads you can inspect, not a ranking of the newest or best. We test a shortlist on your tasks, including the quality of the compressed () version and the license terms.

Same basis for every model: downloads, an , and one active request. Estimated memory covers the model, the cache, and the . Fitting in memory is not a production acceptance test.

Output speeds are unmeasured estimates based on memory , with 4K already in . They are not vendor benchmarks or promised performance. The in each table is a separate example used to size memory. How the estimates work.

Qwen3 4B

Classification, short extraction, a focused assistant.

View on Hugging Face

Start small when the task has a narrow, testable output.

4B
Download at Q4_K_M
2.50 only
32K tokens128K with ; requires testing
License
Apache 2.0

Best fit: Short, well-defined tasks with outputs you can check.

Pros

  • The smallest download in this shortlist, with relatively low memory demand.
  • Supports thinking and non-thinking modes, under the Apache 2.0 license.

Cons

  • Text only; images need a separate model or processing step.
  • Test difficult reasoning and specialist tasks before relying on a model this small.

Scroll sideways for all columns. With a keyboard, focus the table and use the arrow keys.

Hardware planning examples for Qwen3 4B
HardwareWorking contextEstimated memoryOutput speed at 4K
Entry-level workstationNVIDIA RTX 5060 Ti8K tokensInput + output~5.5 50–101 Analytical estimate

View Q4_K_M download files on Hugging Face.

Dense text model

Qwen3 8B

Internal Q&A, extraction, workflow assistance.

View on Hugging Face

A useful baseline to evaluate before buying hardware for a larger model.

Parameters
8.2B
Download at Q4_K_M
5.03 GBWeights only
Native context limit
32K tokens128K with YaRN; requires testing
License
Apache 2.0

Best fit: The first text-model baseline for an internal assistant or extraction workflow.

Pros

  • Fits every hardware tier listed, at the example context sizes.
  • The Apache 2.0 license and adjustable thinking mode make it easy to include in a side-by-side evaluation.

Cons

  • Text only, with no native image understanding.
  • Thinking output and longer context add delay and use memory. Fitting in memory says little about how many users it can serve at once.

Scroll sideways for all columns. With a keyboard, focus the table and use the arrow keys.

Hardware planning examples for Qwen3 8B
HardwareWorking contextEstimated memoryOutput speed at 4K
Entry-level workstationNVIDIA RTX 5060 Ti8K tokensInput + output~7.8 GiB27–55 tok/sAnalytical estimate
AI workstationNVIDIA RTX 509016K tokensInput + output~8.9 GiB111–222 tok/sAnalytical estimate
DGX Spark desktopNVIDIA DGX Spark (GB10)16K tokensInput + output~8.9 GiB16–33 tok/sAnalytical estimate
Apple silicon desktopApple M5 Ultra16K tokensInput + output~8.9 GiB74–149 tok/sAnalytical estimate

View Q4_K_M download files on Hugging Face.

Text and image model

Gemma 3 12B IT

Text and image understanding, document-image review.

View on Hugging Face

Converted to this format by the runtime’s maintainers. Image input also needs the 0.85 GB , and more or larger images add memory and .

Parameters
12B
Download at Q4_K_M
8.15 GBIncludes vision projector
Native context limit
128K tokensActual fit depends on hardware
License
Gemma terms

Best fit: Tasks that need both text and image understanding.

Pros

  • Image input supports visual document review and other multimodal applications.
  • A smaller option than pairing a large text model with a separate vision service.

Cons

  • Gemma’s own license terms need review.
  • The projector, image size, and page count add memory and latency; text-only speed estimates do not apply.

Scroll sideways for all columns. With a keyboard, focus the table and use the arrow keys.

Hardware planning examples for Gemma 3 12B IT
HardwareWorking contextEstimated memoryOutput speed at 4K
AI workstationNVIDIA RTX 50904K tokensInput + output~13.1 GiBBenchmark requiredImage workload varies

View Q4_K_M download files on Hugging Face.

Qwen3 30B-A3B

A compact mixture-of-experts assistant.

View on Hugging Face

All expert weights must fit in memory. reduce the computation per token, not the download size or the memory needed.

Parameters
30.5B total / 3.3B active
Download at Q4_K_M
18.56 GBWeights only
Native context limit
32K tokens128K with YaRN; requires testing
License
Apache 2.0

Best fit: Testing a mixture-of-experts assistant on the 32 GB workstation tier.

Pros

  • Only about 3.3B of its 30.5B parameters are active for each token.
  • This quantization fits a single 32 GB GPU at the example working context.

Cons

  • Every expert’s weights must still be loaded, so it is not a 3B-sized download.
  • Expert routing and the runtime affect speed, so we do not apply the dense-model speed estimate.

Scroll sideways for all columns. With a keyboard, focus the table and use the arrow keys.

Hardware planning examples for Qwen3 30B-A3B
HardwareWorking contextEstimated memoryOutput speed at 4K
AI workstationNVIDIA RTX 509016K tokensInput + output~21.8 GiBBenchmark required varies
Professional AI workstationNVIDIA RTX PRO 6000 Blackwell32K tokensInput + output~23.3 GiBBenchmark requiredExpert routing varies

View Q4_K_M download files on Hugging Face.

Dense text model

Qwen3 32B

More demanding instruction following and analysis.

View on Hugging Face

Compare quality on your own tasks. Parameter count alone is not a quality score.

Parameters
32.8B
Download at Q4_K_M
19.76 GBWeights only
Native context limit
32K tokens128K with YaRN; requires testing
License
Apache 2.0

Best fit: A larger dense text model to compare against the smaller Qwen models.

Pros

  • A dense design avoids mixture-of-experts routing requirements.
  • This quantization fits the 32 GB tier at 16K context.

Cons

  • More memory and computation per token than the smaller dense models.
  • Longer context and concurrent requests reduce capacity, and more parameters do not guarantee better results.

Scroll sideways for all columns. With a keyboard, focus the table and use the arrow keys.

Hardware planning examples for Qwen3 32B
HardwareWorking contextEstimated memoryOutput speed at 4K
AI workstationNVIDIA RTX 509016K tokensInput + output~25.4 GiB30–60 tok/sAnalytical estimate
DGX Spark desktopNVIDIA DGX Spark (GB10)32K tokensInput + output~29.4 GiB4–9 tok/sAnalytical estimate
Apple silicon desktopApple M5 Ultra32K tokensInput + output~29.4 GiB20–40 tok/sAnalytical estimate
Professional AI workstationNVIDIA RTX PRO 6000 Blackwell32K tokensInput + output~29.4 GiB30–60 tok/sAnalytical estimate

View Q4_K_M download files on Hugging Face.

Dense text model

Llama 3.3 70B Instruct

Larger local assistants and complex text tasks.

View on Hugging Face

A community of Meta’s model. Review the license and acceptable-use terms, not just the download format.

Parameters
70B
Download at Q4_K_M
42.52 GBWeights only
Native context limit
128K tokensActual fit depends on hardware
License
Llama 3.3 community license

Best fit: A high-memory text-model evaluation alongside smaller alternatives.

Pros

  • A 70B instruction-tuned candidate with a 128K native context limit.
  • Fits the high-memory single-system configurations listed, at the example context.

Cons

  • Llama’s own license and use terms apply.
  • The quantized weights alone exceed a 32 GB GPU, and full native context needs separate memory sizing.

Scroll sideways for all columns. With a keyboard, focus the table and use the arrow keys.

Hardware planning examples for Llama 3.3 70B Instruct
HardwareWorking contextEstimated memoryOutput speed at 4K
DGX Spark desktopNVIDIA DGX Spark (GB10)32K tokensInput + output~53.6 GiB2–4 tok/sAnalytical estimate
Apple silicon desktopApple M5 Ultra32K tokensInput + output~53.6 GiB9–19 tok/sAnalytical estimate
Professional AI workstationNVIDIA RTX PRO 6000 Blackwell32K tokensInput + output~53.6 GiB14–28 tok/sAnalytical estimate
Multi-GPU system2 × NVIDIA RTX PRO 6000 Blackwell32K tokensInput + output~53.6 GiBBenchmark required layout varies

View Q4_K_M download files on Hugging Face.

Mixture of experts

Qwen3 235B-A22B

Large-model evaluation on a high-memory system.

View on Hugging Face

A capacity example, not our default recommendation. the model across devices adds communication costs that only a complete system benchmark can measure.

Parameters
235B total / 22B active
Download at Q4_K_M
142.15 GBWeights only
Native context limit
32K tokens128K with YaRN; requires testing
License
Apache 2.0

Best fit: A deliberate large-model evaluation after smaller options miss the quality target.

Pros

  • A mixture-of-experts design that uses about 22B active parameters out of 235B.
  • Apache 2.0 gives a familiar licensing basis for evaluation and deployment.

Cons

  • The weights alone exceed one Spark’s memory, before cache and runtime overhead.
  • Sharding adds communication and operating costs, so extra hardware should show a measured quality benefit first.

Scroll sideways for all columns. With a keyboard, focus the table and use the arrow keys.

Hardware planning examples for Qwen3 235B-A22B
HardwareWorking contextEstimated memoryOutput speed at 4K
Multi-GPU system2 × NVIDIA RTX PRO 6000 Blackwell32K tokensInput + output~146.3 GiBBenchmark requiredExpert routing varies

View Q4_K_M download files on Hugging Face.

In these local model profiles, K means 1,024 tokens. Context covers instructions, conversation history, retrieved passages, and the answer. Download sizes use decimal GB; memory estimates use GiB.

Precision guide: FP16, FP8, FP4, and Q4_K_M

Precision is the number of bits used to store each number in a model. Fewer bits can reduce memory use and speed up , but can lower answer quality. The benefit depends on the model, hardware, and working together.

/
Two 16-bit floating-point formats with different numerical ranges. They are the usual higher-precision baseline for checking what changes.
8-bit floating point, in variants such as and . It can reduce weight and activation memory where the hardware and software support it. It is not the same as .
/
4-bit floating point. NVFP4 and use different block-scaling schemes; they are not interchangeable. Scale factors and higher-precision layers add overhead beyond four bits per weight.
The mixed, approximately 4-bit integer quantization used in this guide. It is not FP4 or NVFP4, and it does not make every calculation 4-bit.

This guide’s estimates: Q4_K_M with an FP16 , not FP8 or FP4 benchmarks. A different format, such as , NVFP4, or an conversion, needs its own file-size, memory, speed, and quality checks. Advertised AI are not per second.

NVIDIA Blackwell: FP16/BF16, FP8, and FP4 support. NVFP4 needs a compatible model and runtime; the Q4_K_M examples in this guide use a different format.

Apple silicon: GGUF Q4_K_M through , or separate MLX models. MLX supports 4-bit and 8-bit formats; acceleration depends on the chip, model and runtime.

Technical references: NVIDIA: FP4 and NVFP4; NVIDIA: RTX runtime support; Hugging Face: GGUF formats; MLX quantization formats.

Beyond memory fit: expert models, licenses, and disk space

: Qwen3 30B-A3B activates about 3.3B per token, but its full 18.56 download still has to fit in memory. Do not size it like a 3B dense model.

Licensing: “” does not mean unrestricted. Qwen uses Apache 2.0 for these ; Gemma and Llama have their own terms. Check what each license allows for use and redistribution, and what it requires of fine-tuned versions.

Disk is not memory: A 20 GB model download does not mean a machine with 20 GB of memory is enough. Keep disk space for at least two model releases, images, document indexes, source files, and independent backups.

Hosted AI reference

How does local AI compare with ChatGPT, Codex, and Claude?

ChatGPT, Codex, and Claude do not have one fixed speed. The model you select, its reasoning mode, the tools it uses, and how busy the service is all change the experience. The table below compares specific models through their instead, which gives a more useful reference.

Reference snapshot: . Speeds are approximate figures from Artificial Analysis for the listed provider and settings, with 10,000 input . specifications and prices are not the same as ChatGPT, Codex, Claude, or Claude Code subscription limits.

Not a like-for-like benchmark. The hosted figures are measured API observations; the local figures above are unmeasured planning estimates at 4K occupied . The comparison does not show equal speed, answer quality, or coding ability.

Scroll sideways for all columns. With a keyboard, focus the table and use the arrow keys.

Standard API prices in US dollars per million tokens. Speeds are rounded, and time to first answer includes thinking time. In this hosted table, K = 1,000 and M = 1,000,000.
Model and API providerOutput speedFirst answerContext / API input / output
GPT-6 AstraOpenAI, High reasoning~53 GPT-6 Astra benchmark~33.3 secBefore answer starts1.05M / 128K922K $10.00 / $50.00$1.00
GPT-5.6 SolOpenAI, High reasoning~55 tok/sGPT-5.6 Sol benchmark~9.3 secBefore answer starts1.05M / 128K922K max input$4.00 / $20.00$0.40 cached input
GPT-5.6 TerraOpenAI, High reasoning~92 tok/sGPT-5.6 Terra benchmark~2.5 secBefore answer starts1.05M / 128K922K max input$2.00 / $12.00$0.20 cached input
GPT-5.6 LunaOpenAI, Non-reasoning~109 tok/sGPT-5.6 Luna benchmark~0.9 secBefore answer starts1.05M / 128K922K max input$0.20 / $1.20$0.02 cached input
Claude Fable 5.1Anthropic, High, default fallback~51 tok/sClaude Fable 5.1 benchmark~10.9 secBefore answer starts1M / 128KInput and output share context$10.00 / $50.00$0.25 cached input
Claude Opus 5Anthropic, High, adaptive thinking~49 tok/sClaude Opus 5 benchmark~8.6 secBefore answer starts1M / 128KInput and output share context$5.00 / $25.00$0.50 cached input
Claude Sonnet 5Anthropic, High, adaptive thinking~58 tok/sClaude Sonnet 5 benchmark~2.3 secBefore answer starts1M / 128KInput and output share context$2.00 / $10.00$0.20 cached input
Claude Haiku 4.5Anthropic, Non-reasoning~81 tok/sClaude Haiku 4.5 benchmark~0.7 secBefore answer starts200K / 64KInput and output share context$1.00 / $5.00$0.10 cached input

Artificial Analysis reports a rolling 72-hour and normalizes to OpenAI tokens; prices and context limits use each model’s native tokens. The figures are a dated snapshot, not a service guarantee, and were not measured in the consumer apps. Benchmark methodology.

How to read the comparison: speed, context, pricing, and deployment

Fast output does not mean a fast completed task

At 50 tokens per second, 500 output tokens take about 10 seconds once generation starts. Prompt processing, thinking, retrieval, and tool calls add time, and an that edits code and runs tests may make many model calls. Measure the time to finish a task, not just the speed of the text stream.

API limits are not app limits

The table lists each model’s maximum, not space guaranteed for your documents. Instructions, conversation history, retrieved passages, and the answer all share that budget. Chat products and coding may reserve space, shorten history, or apply different limits, and which models you can use varies by plan and environment. Codex model options; Claude model specifications.

Per-token prices are only part of the cost

For the OpenAI models listed, when a prompt exceeds 272K input tokens, the entire request is charged at 2× the input and rates and 1.5× the output rate. Sol’s listed promotional pricing is available at least through November 21, 2026. The Claude models listed use standard pricing across their full context windows. Cache writes, faster service tiers, batch processing, tools, and regional options can change the bill. Prices are subject to change. OpenAI pricing; Anthropic pricing.

Running models locally avoids per-token charges, but equipment, electricity, software operations, storage, and replacement still cost money. Compare the cost of each accepted result at your expected workload, not the hardware price against a chat subscription.

Hosted models do not publish sizing details

These hosted models have no downloadable and no public sizing specifications for an build. We do not guess their parameter counts, , or / formats. The models above are separate choices, not local editions of ChatGPT or Claude.

Document search (RAG) and tools belong to the application

The listed API models accept text and images and produce text. Voice, search, file handling, and code execution inside a product rely on additional services. Hosted and local models can both sit behind document retrieval and approved tools, but the application still has to provide access control, citations, validation, and action confirmation.

A hosted API sends each request to its provider, so review that provider’s retention, region, and contract terms. An on-premises design keeps processing on site only when retrieval, , logging, tools, and other dependencies stay inside the same boundary.

Choose by task quality and deployment requirements

A smaller local model may suit a narrow task even when a hosted model is stronger overall. We test both with the same examples, acceptance criteria, and simultaneous load. Decide which model meets the requirement before choosing hardware for it.

The application around the model

Document search, images, speech, and agents

The same machine can support more than a chat model. Each extra workload uses memory, storage, or processing time, so we size the whole application, not just the model file.

Search your documents with RAG

finds the relevant passages in your documents and gives them to the model with the question. The document library itself does not have to fit in the or .

Parsing, , , retrieval, and any run locally too. We apply source permissions before retrieval, keep citations, and handle document updates and deletions.

Storage example: 100,000 × 768 dimensions × 4-byte floats = 307.2 MB of raw . A million chunks = 3.07 . Text, metadata, index overhead, originals, and backups are extra. This is arithmetic, not a claim about how many documents a system can hold.

Read images and scanned documents

A vision-language model can interpret an image; OCR extracts text for search and structured processing. Choose based on the job, not the label “.”

Gemma’s is an additional file. Image resolution, page count, and batch size change memory and . Text token speeds do not predict pages per minute.

Vision runtime requirements

Transcribe speech or run scheduled jobs

Local speech recognition, extraction, and classification can run as separate services. Non-urgent work can wait in a queue so it does not compete with an interactive assistant.

We measure speech in minutes of audio processed per minute, stating the language and noise conditions. Training, image generation, and need separate sizing; the tiers above cover only.

Connect an agent to internal software

An needs approved tools, identity, scoped permissions, action review, and recovery. A model’s ability to call tools is not permission to change a record.

Acceptance tests include multi-step delays and tool failures. Credentials stay outside the model, and we check whether a simpler workflow would do the job better.

APIs and agent integrations
Long and are different tools

For Qwen3 8B, the for one request is about 1.13 at 8K or 4.5 GiB at 32K, before model and overhead. Four 32K requests running at once can need about 18 GiB of cache alone. RAG keeps prompts short by retrieving a focused set of passages instead of loading the whole library into every prompt.

How cache memory works

What the estimates mean

No hardware benchmarks were run for this guide. The planning calculations use published specifications and exact file sizes from pinned model versions.

Memory and context

Estimated memory is the model files, plus one key/value , plus a stated allowance for the . Each tier keeps spare memory outside that estimate. Real runtime reservations, image processing, and can use more.

is a conservative pilot setting, not the largest possible window. Gemma’s 4K text-first example uses a separate allowance because its hybrid attention does not follow the simple full-cache calculation.

Output speed, not total response time

The dense-model range assumes 35–70% of advertised memory , divided by weight bytes plus cache bytes at 4,096 occupied . That efficiency range is an uncalibrated sensitivity assumption, meaning an assumed spread rather than a measured confidence interval.

It assumes the model is already loaded and fully accelerated, with one active request and no . It leaves out model loading, prompt processing, retrieval, tools, and time to first token. Compute limits, kernels, and memory traffic can put real results outside the range. We give no number for expert models, vision, or layouts, where this shortcut does not hold.

What we verify before purchase

We pin the model, , runtime version, driver, , and hardware configuration. We test representative 2K, 8K, and longer prompts, then the intended simultaneous load.

We record answer quality, time to first token, output per request, total , peak memory, power, and failure recovery. Warm runs are repeated, and cold starts are tested separately.

Benchmarking reference

Showing that a model fits is not a claim of customer deployment, compatibility certification, or vendor partnership.

Discuss hardware and deployment

What you get

A deployment plan

A map of where your data flows, a hardware specification, a model shortlist, a licensing review, a cost outline, and written acceptance criteria. We assess your existing equipment before recommending any purchase.

A working application

Local model serving, an interface or , and connections to approved internal sources. Retrieval, document processing, permissions, and review steps are included where the task needs them.

An operating handoff

Deployment configuration, an inventory of versions and dependencies, evaluation results, administrator training, and step-by-step runbooks for updates, backups, recovery, and support.

How on-premises AI is designed and deployed

We lead the technical work alongside your business and IT owners. Your team explains the task, approves access, reviews the pilot, and accepts the release. We agree on scope and cost before each stage, and a pilot is not a commitment to buy hardware or launch.

  1. Choose the first workflow

    We walk through one real task with your team, such as finding answers in internal manuals or extracting fields from documents. We note who does the work, the exceptions, and what goes wrong today. Together we agree on what a good result looks like, including when the AI should ask for help or take no action.

    Result: A focused scope, a named owner, and measurable acceptance criteria.

  2. Map the data and design the application

    We trace where documents, prompts, outputs, logs, and backups go. With your IT and security owners, we decide who can access each source and whether the network is restricted or (fully offline). We design the interface, local processing, storage, and identity connections, and check whether retrieval from your documents is enough before considering .

    Result: An architecture and data-flow map, with access rules and a pilot test plan.

  3. Prove the approach on a local pilot

    We use fictional or redacted examples until access, storage, logging, and network controls are approved. We compare suitable licensed models on representative tasks, then test with separate evaluation examples, measuring accuracy, response time, concurrent use, and memory. If the results do not justify a build, we revise the approach or stop.

    Result: Pilot findings, a model recommendation, and measured hardware requirements.

  4. Prepare the production environment

    After budget approval, we reuse suitable equipment or help arrange the agreed hardware; the scope sets who buys and installs it. We configure storage, encryption, identity, backups, and network restrictions, and install approved, versioned models and dependencies through your permitted import process. We check power, cooling, capacity, and recovery arrangements before production data is introduced.

    Result: A configured environment, verified dependencies, and an installation record.

  5. Connect the data and build the workflow

    We connect approved sources and prepare documents for local search or extraction. We build the application around your existing permissions and handle source updates, deletions, and revoked access. When an can take actions, we enforce permissions in application code, keep credentials outside the model, and require the agreed review before any change.

    Result: A working application connected to approved internal systems.

  6. Test the complete system

    Representative users from your team test ordinary tasks, missing information, and difficult cases. We check forbidden access attempts, outbound traffic, load, restarts, and offline operation where required, and rehearse backup restoration and . We document failures and resolve anything that blocks release under the acceptance criteria.

    Result: An acceptance report with measured results, known limits, and a release decision.

  7. Launch with a small user group

    With your approval, we release to a defined group before expanding access. We train users and administrators, monitor the agreed measures locally, and keep rollback available. We hand over the configuration, operating instructions, incident contacts, and a clear split of responsibility for hardware, software, and model behavior.

    Result: An accepted deployment and a team prepared to operate it.

  8. Maintain and improve it

    Your team or an agreed support arrangement owns this work; remote access is not assumed. Those responsible review quality, capacity, source freshness, and access as usage changes. They stage approved patches and model updates, then repeat quality and data-boundary tests before release. They also keep backups and recovery checks current.

    Result: A maintenance schedule, named responsibilities, and controlled future releases.

Example applications for sensitive information

These are illustrative build options, not claims about completed customer deployments. We start with one task and expand only when it earns its place.

An internal knowledge assistant

Staff ask questions of approved policies, technical manuals, or research. Answers cite their sources, respect each user’s access, and say when the documents do not support an answer. The design handles source updates, deletions, and revoked access.

Document review and extraction

The application reads contracts, case files, or internal reports locally. It extracts fields, compares versions, and sends uncertain results to a reviewer instead of silently updating the record.

Agents for internal software

An assistant prepares a task or record change through approved . Credentials stay outside the model, actions are limited, and a person confirms consequential changes.

What stays inside the boundary

When no data may leave your site, prompts and documents are only part of the picture. We account for every component that handles your data and verify the agreed configuration before acceptance.

Processing and storage

Model , , , and any agreed run locally. Outputs, attachments, and retrieval indexes are stored on site. No external inference endpoint, hosted , or automatic cloud fallback is allowed.

Logs, backups, and dependencies

Logs, caches, backups, monitoring, and identity services stay inside the agreed boundary. We disable , crash uploads, external analytics, and automatic downloads, and we enforce network restrictions instead of relying on application settings alone.

People and support

We define with you who can view records, export files, use removable media, or administer the deployment. A strict on-site scope rules out remote viewing and off-site backups. Support then happens on site or through procedures your staff run, not through unapproved log or screen sharing.

On-premises, restricted, or air-gapped?

We document which hardware and network services may process your data. A private cloud still uses off-site infrastructure, so it is not an on-site deployment and has different requirements.

On-premises

The application and models run on hardware at your site, and connections to internal systems remain available. When data must stay on site, we also restrict outbound traffic and review user devices, identity, and backup locations.

Restricted network

Only documented connections are allowed, and we review any proposed external service on its own. If an exception processes business data off site, we do not describe the deployment as one where no data leaves your site.

Air-gapped

The environment has no network connection to outside systems. Models and updates arrive through approved offline import procedures, with integrity checks, dependency review, and a plan. Whether this works depends on the complete hardware and software stack.

Deployment options

You do not need to have chosen a model or bought servers. We can assess first, deliver one application, or build a shared platform in stages.

Readiness and architecture assessment

We work out feasibility, the data boundary, hardware options, and total operating costs. You receive a pilot plan or a clear recommendation not to proceed. Implementation is a separate decision.

Focused deployment

We take one workflow from a local pilot to an accepted production release, including installation, integrations, controls, documentation, and staff handoff.

Internal AI platform

We build shared model serving, identity, monitoring, and release controls for several applications. Capacity and use cases grow through separately approved stages, with an ongoing maintenance agreement if you need one.

Scope, cost, and ownership

What we need from you

  • A business owner and an IT or security counterpart; a non-sensitive task description is enough to begin
  • An agreed site and data boundary, including permitted access and support arrangements
  • Approved representative examples and test users, reviewed inside your environment when required

What affects cost

  • Model size, length, concurrent users, , and quality requirements
  • Hardware procurement, power, cooling, storage, redundancy, and support coverage
  • Source preparation, integrations, identity, licensing, and offline update requirements
  • Security review, acceptance testing, staff training, and ongoing maintenance

Technical scope

  • Compatible local model and , with deployment and commercial-use terms checked for each model
  • Least-privilege identity, encrypted transport and storage, key ownership, and retention controls
  • Local processing dependencies, with outbound network access blocked or explicitly controlled
  • Versioned deployments, approved update imports, health monitoring, tested backups, and

Support and maintenance

We lead application architecture and delivery, coordinating with your IT team and with hardware or security specialists where needed. The agreement assigns responsibility for equipment, networking, model quality, updates, incidents, and support hours. We do not assume unattended remote access or 24/7 coverage.

Common questions

How much hardware do we need for a local AI model?

It depends mainly on the model, its length, and how many requests run at the same time. Our guide on this page compares equipment budgets from a small 16 pilot to high-memory workstations and systems, with example model downloads and memory calculations. Fitting in memory is a planning check, not proof of useful speed or answer quality, so we benchmark the complete application before you buy. Implementation and support are priced separately.

Does a RAG document library have to fit in GPU memory?

No. stores searchable documents and an index separately, then passes selected passages to the model. still needs room for model , , and any embedding or reranking models running there. We also size document ingestion, access controls, storage, backups, and simultaneous requests as part of the application.

Can you build AI where no data leaves the building?

Yes. We can scope a fully local deployment with no external processing. The commitment covers a defined application and operating boundary, including logs, backups, identity, user devices, and support. Installing a local model alone does not prove it, so we test that configuration before acceptance and after any material change.

Can we build a DGX Spark cluster?

Yes. Our guide compares two-, three-, and four-node clusters that follow NVIDIA’s documented connection setups, plus an eight-node option that needs a custom engineering review. Node prices exclude networking and deployment. Combined memory is not automatically one usable pool, so we test the model, , network layout, and application together.

Can a local model meet our requirements?

It depends on the task, so we test before recommending a production build. We compare task accuracy, source grounding, failure handling, speed, and capacity against your acceptance criteria. A smaller local model may fit a narrow task; other workloads may need more hardware, a different approach, or a decision not to proceed.

Do we need to train a model from scratch?

Usually not. We start with a suitable licensed model, approved retrieval sources, and a focused application. is an option only when measured task failures justify it and the data rights, compute, and evaluation are in place.

Can we use our existing servers?

Yes, if they meet the workload’s requirements. We review processor and GPU support, memory, storage, power, cooling, backup power (UPS), expected concurrent use, and availability requirements, then benchmark the intended workload before you commit to equipment. The scope sets who buys and installs any hardware.

Will on-premises AI cost less than a hosted service?

Not automatically. A fair comparison includes hardware, utilization, power, licenses, IT time, evaluation, support, and replacement costs. Where external processing is permitted, a hosted service may be the better choice. We do not promise savings or a payback period without a workload-specific calculation.

Does on-premises mean compliant or immune to data leaks?

No. Location is one control, not a certification or a guarantee against misuse. Access, retention, endpoint security, vulnerabilities, exports, and human decisions still matter. Your security and compliance owners approve the required controls; we provide the agreed technical evidence.

Can support and updates work without internet access?

Yes, if the chosen software stack supports offline operation. The plan covers approved model and software imports, integrity checks, staging tests, , and maintenance done on site or by your staff. Remote sessions or off-site diagnostic exports need explicit approval and may conflict with a strict on-site boundary.

What should we send in the initial inquiry?

Only a non-sensitive description of the task, your deployment restrictions, the approximate number of users, and your known infrastructure. This public website and its contact form are not an environment. Do not send confidential files, records, credentials, or network diagrams; we agree on a suitable review process first.

Guides and resources

See also

Have a deployment like this in mind?

Discuss an on-premises deployment