{"rewrite":{"id":"r_32a7dc7b1d487b38d904cb00","clusterId":"c_af0d69687957965af8d9499e","slug":"nvidia-releases-pair-a-free-tool-that-routes-ai-tasks-across-your-local-network","model":"deepseek-v4-flash","headline":"NVIDIA Releases PAIR, a Free Tool That Routes AI Tasks Across Your Local Network","summary":"NVIDIA has released PAIR (Personal AI Router), a free virtual inference router that detects compatible PCs on a local network and assigns AI inference jobs to them. It works with Ollama and LM Studio and needs no new API. In a demo running five subagents with Qwen 3.6-35B-A3B, a three-machine cluster averaged 8 minutes 48 seconds versus 18 minutes on a single RTX Spark PC.","whyItMatters":"PAIR turns idle machines on a home or office network into a pooled inference cluster, addressing the GPU bottleneck that multi-agent workflows create on a single machine.","webCardHtml":"\u003cp\u003ePAIR\u0026#39;s proxy is compatible with Ollama and LM Studio, so it requires no new API. When no jobs are queued, it powers off or hibernates idle nodes.\u003c/p\u003e\u003cp\u003eThe demo paired PAIR with Hermes Desktop, a GUI for the Hermes Agent, and Ollama to route five subagents. With Alibaba\u0026#39;s Qwen 3.6-35B-A3B, a cluster of an RTX Spark PC, a DGX Spark workstation, and an RTX 5090 PC averaged 8 minutes 48 seconds, against 18 minutes on a single RTX Spark PC.\u003c/p\u003e\u003cp\u003eRequirements cover Windows 11, DGX OS, Ubuntu 14.04, or macOS Tahoe, a GeForce RTX 20 series or newer, DGX Spark, or Mac M4 or newer, 8GB of memory, and 20GB of free space.\u003c/p\u003e","blueskyPost":"NVIDIA released PAIR, a free virtual inference router that finds compatible PCs on your local network and routes AI inference jobs to them. It works with Ollama and LM Studio. A three-machine cluster finished a five-subagent workload in 8:48, versus 18 minutes on one PC.","twitterPost":"NVIDIA released PAIR, a free virtual inference router that detects compatible PCs on a local network and assigns AI inference jobs to them. Works with Ollama and LM Studio. A three-machine cluster ran a five-subagent workload in 8:48 vs 18 minutes on a single PC.","threadsPost":null,"newsletterBlurb":"NVIDIA has released PAIR, a free virtual inference router that pools idle PCs on a local network into a shared inference cluster. It works with Ollama and LM Studio, and a demo showed a three-machine cluster finishing a five-subagent workload in under nine minutes against 18 on a single machine.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260904-nvidia-personal-ai-router/\",\"title\":\"NVIDIA Releases PAIR, a Free Tool That Routes Heavy AI Processing to PCs on the Same Network\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":5367,"outputTokens":4628,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":null,"createdAt":"2026-09-04T04:58:24.000Z","publishedAt":"2026-09-04T05:01:41.000Z","updatedAt":"2026-09-04T05:01:41.000Z"},"cluster":{"id":"c_af0d69687957965af8d9499e","canonicalTitle":"NVIDIAが重たいAI処理を同一ネットワーク内のPCに割り振れる無料ツール「PAIR」を公開","representativeArticleId":"a_72aa46aee72336704df0413b","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[\"PAIR\"],\"studios\":[],\"people\":[],\"type\":\"announcement\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-09-04T04:00:00.000Z","lastSeenAt":"2026-09-04T04:00:00.000Z","updatedAt":"2026-09-04T05:01:40.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20260904-nvidia-personal-ai-router/","title":"NVIDIAが重たいAI処理を同一ネットワーク内のPCに割り振れる無料ツール「PAIR」を公開"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":["PAIR"],"studios":[],"people":[],"type":"announcement","domain":"other","is_roundup":false},"keyFacts":null}
