# Offline Base llms.txt

This is a single Markdown export of Offline Base site content, product facts, route map, and source documents. It is intended for copy-all use and agent traversal.

The export intentionally keeps visual components, animation details, and binary assets out of the body while preserving the text, product facts, model data, and source Markdown that explain the product.

## Site Map

Canonical routes and short descriptions:

- `/` - **Home**: Offline Base overview and Base Drive purchase entry point.
- `/base-drive` - **Base Drive**: The $59 USB drive that sets up private offline AI on a laptop.
- `/how-it-works` - **How It Works**: How Base Drive setup, consent, and local-only use work.
- `/use-cases` - **Use Cases**: The full catalog of private, local AI use cases.
- `/about` - **About**: Why Offline Base exists and the promises behind the product.
- `/store` - **Shop**: Preview the Base Drive lineup ahead of launch.
- `/private` - **Privacy**: The privacy case for local AI on hardware the buyer owns.
- `/privacy` - **Base App Privacy**: The privacy policy for the Base iPhone app from Offline Base LLC.
- `/support` - **Base App Support**: Setup help, troubleshooting, and contact information for the Base iPhone app.
- `/the-box` - **Base Router**: The upcoming always-on private AI product for a home or office.
- `/integrations` - **Integrations**: Planned local memory connections for assistants and MCP tools.
- `/try` - **Try It**: A live chat with a small model running on our box, the Base Router hardware.
- `/open` - **Open Models**: A free chat with curated open-weight models, favoring zero-data-retention providers.
- `/benchmarks` - **Benchmarks**: Model comparisons and honest product performance expectations.
- `/professionals` - **For Professionals**: Private AI for sensitive and regulated professional work.
- `/independent` - **Independence**: The ownership case for AI that is not rented from a cloud provider.
- `/sustainable` - **Sustainability**: The energy case for answering everyday AI tasks locally.
- `/build-your-drive` - **Build Your Own Drive**: A preview model picker, drive sizer, and price estimator.
- `/build` - **Site Builder Demo**: A permanent redirect to the site-builder tab on the Try page.
- `/cart` - **Cart**: Shopify cart and secure checkout handoff.
- `/links` - **Links**: A compact directory of Offline Base products, demos, and profiles.
- `/llms.txt` - **LLMs Text**: The public Markdown export for copying and agent traversal.

## Home

Route: `/`

### Your own AI. Private by default.

Base Drive sets up private AI on the laptop you already own in minutes. It is coming soon for $59, works offline, has no subscription, and keeps conversations on the owner's machine. No product is orderable yet.

Compatibility: Apple Silicon Macs and Windows 11 PCs.

### NVIDIA Inception

Offline Base is a member of NVIDIA Inception.

### Two words worth knowing

- **Local AI:** The model runs on the owner's own computer instead of a company's server. What they type never leaves the machine, and it keeps answering with the wifi off.
- **Open source AI:** The model's weights are published for anyone to download, inspect, and run. Once someone has a copy, no one can retire it, throttle it, or change the terms on them.

### Privacy, price, and control

- **Privacy:** There is nothing to upload, so there is nothing to leak. Drafts, files, and questions stay on the owner's machine. If a feature ever needs outside help, they see what would leave and decide.
- **Price:** Base Drive is $59, once. No monthly bill, no token meter, no account to keep alive. It keeps working whether or not the owner keeps paying.
- **Control:** The owner holds the model, so they decide when it changes. Nothing updates behind their back, and no version they rely on gets deprecated out from under them.

### Try an open model, then bring it home

Chat free with 10 open-weight models in the browser at /open, routed only to providers that keep no record of it. Base Drive sets the same idea up on the owner's own laptop in minutes for $59.

Base Router is in development: the same private AI on an always-on box for the whole home. It is not available to order.

## About

Route: `/about`

### We think you should own your AI.

Offline Base is a small company with a simple position: AI has become part of daily life, so it should live on hardware you own, under rules you set.

Open models became good enough and small enough to run on the laptop people already own, but most people still meet AI as a subscription, a meter, an account on someone else's computer, and their questions in someone else's logs.

Offline Base closes that gap with products that put the models in the buyer's hands: a drive that sets up a laptop, a box that serves a home, and a memory layer that stays theirs no matter which assistant they use.

### Four promises

- **Private:** Your data stays under your control. Nothing leaves the box without your explicit approval, and you can turn outside access off entirely.
- **Local:** Runs on the box, works offline, and does not require another account or a usage meter.
- **Fine-tuned:** Not a generic chatbot. Prepared around your documents, routines, and use case.
- **Packaged:** It just works. Plug it in, open it from your devices, and start using it. No technical setup.

### Product forms

- **Base Drive:** first product. A $59 USB drive that sets up private, offline AI on a laptop in minutes.
- **Base Router:** next up. The same private AI moved onto dedicated, always-on hardware for the whole home.
- **Base Connect:** coming. A connector that brings local memory to ChatGPT and Claude, with Base deciding what they see.

### How we work

- We'd rather under-promise. A small local model is not the largest cloud system, and the product says so before checkout.
- Nothing happens silently. No telemetry, quiet updates, or hidden cloud calls. When outside help is needed, the user approves it first and a record is kept.
- Workshop video: Learn to run LLMs locally at CUNY, embedded from YouTube.

## Base Drive

Route: `/base-drive`

### Your own AI. On a drive.

A USB drive that sets up private, offline AI on the laptop the buyer already owns in a few minutes. No cloud account. No API bill. No setup headache.

Works with Apple Silicon Macs and Windows 11.

### What is on the drive

The drive carries everything; the laptop does the thinking. The installer copies AI to the machine so it runs at full speed, even with the drive unplugged.

- **Install Offline Base.app** (Application, Mac and Windows): One-click setup that copies the models, verifies them, and opens the chat app. No terminal, accounts, or manual model wrangling.
- **models** (Folder, 2 items, 3.9 GB): Two GGUF models ship today: a compact model for lighter machines and a stronger everyday model. More models are on the way.
  - everyday-model.gguf, 3.0 GB
  - compact-model.gguf, 0.9 GB
- **ui** (Folder, chat app): Open WebUI in the browser, running offline against the model on the user's own machine.
- **memory.md** (Markdown, coming soon): A plain file that holds what the AI remembers. The user can open, edit, or delete it.
- **docs** (Folder, readable offline): Offline documentation for running, owning, and trusting local AI.
- **VERSION.txt** (Plain text): Updates arrive as a fresh drive, not a background download. Nothing updates itself silently.

### Will it run on your laptop?

A drive cannot make a weak laptop fast, so the installer asks what hardware is available and recommends the model that fits.

| Device | Verdict |
| --- | --- |
| MacBook with Apple Silicon (M1 or newer) | Great |
| Windows 11 PC with an NVIDIA GPU | Great |
| Windows 11 laptop, 16 GB RAM | Good |
| Windows 11 laptop, 8 GB RAM | Small model only |
| Intel Mac, Chromebook, phones | Not supported |
| Linux | On the roadmap |

This is a small local model, not ChatGPT. It is a private writing assistant, study buddy, and document helper the buyer owns.

### Why trust a USB drive?

- **Anyone can check nothing was tampered with:** The installer is open source and every file on the drive is listed. Each drive ships with SHA-256 checksums, and setup verifies them before copying anything.
- **Your laptop may show a warning:** The installer is not signed with Apple or Microsoft yet. Signing and notarization are still in progress, and every file can be inspected before it runs.
- **Nothing phones home:** No account, no telemetry, no usage meter. The app talks to the model on localhost, and it keeps working after Wi-Fi is turned off.

### Drive tiers

| Tier | Status | Price | Storage | Models | Best for |
| --- | --- | --- | --- | --- | --- |
| Base Drive Lite | Preview only | $39 | 64 GB USB-C | One small everyday model, tuned for older laptops | Weak laptops and first-timers who just want to see it work. |
| Base Drive | Coming soon | $59 | 128 GB USB-C + USB-A | Two everyday models today, sized for your laptop. More on the way | Most people. Students, writers, and anyone done renting AI. |
| Base Drive Pro | Preview only | $79 | 256 GB USB-C + USB-A | Larger models plus a coding-focused setup | Developers who want local inference without the API bill. |

## Base Router

Route: `/the-box`

### One box. One promise: it just works.

Base Router is in development. It is designed as a finished appliance with the software and AI already loaded, ready to join a network without a keyboard, monitor, or technical setup.

Base Router is the step after Base Drive: the same private AI, moved onto dedicated, always-on hardware.

### Base Router

- Status: Upcoming; not available to order
- Planned price: $400
- Tagline: The full package, fine-tuned to you.
- Hardware: Jetson Orin Nano Super, a small NVIDIA computer (8 GB)
- Memory: 8 GB
- AI included: Larger everyday AI planned for documents and professional work
- Pace: A steady conversational pace
- Power: 7 to 25 watts
- Best for: A capable everyday assistant, document work over your own files, and field-specific setup.
- Straight talk: This is designed to be a strong box, but it will still be a desktop product. For the hardest questions, it may ask permission to use outside help.

### What's inside

No mystery hardware: a shell, a compute board, and storage where the models and documents live as ordinary files. There is no microphone and no camera.

- The AI is files on the storage layer, yours to inspect.
- Your documents never leave that same layer.
- Nothing inside listens, watches, or phones home.

### Speed

Bigger models read better but generate slower. The demo previews the planned Base Router hardware so visitors can understand its target pace.

- **IBM Granite 4.1 8B:** Tuned for documents and business tasks. Base Router pace: 15 tokens per second.
- **Qwen 3.5 9B:** Strong reasoning. The heavier, slower pick. Base Router pace: 12 tokens per second.

### Honest expectations

- It is not ChatGPT. The included AI is useful but will not match the largest cloud systems on the hardest questions.
- Speed is practical, not instant. Long answers and large documents take more time.
- Base Router is planned with active cooling.
- Updates can ship on a USB stick. Updates are checked before install, and the previous version is kept for rollback.

## How It Works

Route: `/how-it-works`

### Plug it in. Open the app. Start using private AI.

Offline Base makes local AI setup disappear, then stays clear about what happens after.

1. **Plug it in:** Connect Base Drive to a supported Mac or Windows PC. Nothing runs automatically.
2. **Run one installer:** The installer verifies the files, picks a model that fits, and sets up the local chat app.
3. **Use AI offline:** Unplug the drive after setup. The model and conversations stay on the owner's computer.

### Consent

Base Router is upcoming and not available to order. Its planned consent flow handles questions locally first, asks before outside help, records each approval, and includes a local-only mode.

- A prompt before a single word leaves the box.
- A written record of every approval.
- A local-only mode that turns outside help off entirely.

### Remote access

Remote access is a planned, optional Base Router feature. When enabled, the connection will be encrypted end to end. Some outside coordination may establish the secure link, so the feature will start off and require an explicit choice.

Base Drive is the available product today and works offline on the owner's laptop after setup.

## Privacy

Route: `/private`

### Your thoughts aren't training data.

Eyebrow: The privacy case

Every question typed into a cloud chatbot is handed to someone else's system. Offline Base runs on a box the buyer owns, so what they ask stays with them.

### Points

- Nothing is logged by Offline Base: no account, no telemetry, no cloud to log to.
- If a task needs outside help, the user gets a consent prompt first and a written record after.
- The built-in app shows what happened. No black box, fine print, or hidden outside vendors.

### The receipts

- May 2025: a federal judge in the New York Times copyright case ordered OpenAI to preserve ChatGPT output logs it would otherwise delete, including chats users had deleted. The blanket hold ended that September, but the captured data (except logs from the EEA, Switzerland, and the UK) stays held for the litigation. (Source: Malwarebytes Labs)
- January 2026: a federal judge affirmed an order for OpenAI to produce 20 million de-identified consumer ChatGPT conversations, full prompts and full answers, to the publishers suing it. (Source: Robinson+Cole, Data Privacy + Cybersecurity Insider)
- ChatGPT's consumer 'Improve the model for everyone' training setting is on by default, and opting out only works going forward; data already used in training is not removed. (Source: Tom's Guide via Yahoo Tech)
- January 2025: security researchers found DeepSeek had left a database publicly reachable with no password, exposing over a million lines of logs including users' chat histories and API secrets. (Source: Wiz Research)

> The only AI you can fully trust with a secret is one that physically cannot share it.

### The honest part

Once outside help or remote access is allowed, 'data never leaves the box' becomes conditional. Those choices are visible, opt-in, and easy to turn off.

## Independence

Route: `/independent`

### Stop renting your mind from five companies.

Eyebrow: The independence case

Cloud AI is a meter the buyer does not control, on terms they did not write. Offline Base is bought once and owned outright.

### Points

- In the past year, cloud models were switched off by government order, pulled from paid plans, and swapped without warning. The box has no account to suspend and no switch anyone else can flip.
- No usage caps or surprise invoice. The cost does not move with daily questions.
- Outages, rate limits, and policy changes elsewhere do not reach the box.

### The receipts

- June 2026: a US government export-control directive forced Anthropic to suspend Claude Fable 5 for every customer, including its own employees, three days after release. Access returned in early July. (Source: Anthropic)
- July 2026: days after restoring it, Anthropic announced Fable 5 would leave Pro and Max subscriptions anyway, moving to pay-per-use, citing demand that was 'very high, and difficult to predict.' (Source: BleepingComputer)
- August 2025: OpenAI removed GPT-4o and other older models from ChatGPT for all but Pro users, without advance warning, the day GPT-5 launched, then brought it back for paid users after backlash. (Source: TechCrunch)
- February 2025: Humane's $699 AI Pin lost calling, messaging, and all AI features ten days after HP bought Humane's assets. Refunds covered only devices shipped in the previous 90 days. (Source: TechCrunch)

> Don't rent your thinking from a company that can change the terms while you sleep.

### The honest part

Independence means the owner owns upkeep too. The box will not match the largest cloud systems on the hardest tasks, so outside help exists only on the user's terms.

## Sustainability

Route: `/sustainable`

### AI that doesn't need a power plant.

Eyebrow: The sustainability case

Cloud queries run in data centers. Offline Base answers from a box on the desk that sips power.

### Points

- Base Router draws between 7 and 25 watts.
- The question is answered in the room instead of round-tripping to a data center.
- One box used for years beats a constantly refreshed fleet of unseen cloud hardware.

### The receipts

- The IEA counted about 415 TWh of data-center electricity use in 2024, roughly 1.5% of the world's total, and projects it to more than double to around 945 TWh by 2030, with AI the most important driver. (Source: IEA, Energy and AI)
- Google's water consumption rose 28% year over year to 8.1 billion gallons in 2024, while its data-center electricity use grew 27% in the same year and has doubled over four years. (Source: RCR Wireless, from Google's 2025 Environmental Report)
- Google estimates a median Gemini text prompt uses 0.24 Wh of energy and about five drops of cooling water. That is Google's own math, not independently verified, and covers only median text prompts. (Source: Google Cloud Blog)
- NVIDIA benchmarked the module inside Base Router at about 43 tokens per second on a 3B model in its full 25-watt mode; by spec-sheet arithmetic, not a lab measurement, a 200-token answer is on the order of 0.03 Wh. (Source: NVIDIA Developer Blog)

> The greenest query is the one that never leaves the building.

### The honest part

A desktop box is not carbon-free. The narrower claim is that everyday local tasks avoid data-center overhead and keep working for years.

## Integrations

Route: `/integrations`

### Bring your memory to the AI you already use.

Base Connect plugs Base into ChatGPT and Claude as a connector. They can ask Base for the few facts a task needs and nothing more. Memory stays on hardware the user owns. Coming as a free software update: Base Drive first, then Base Router.

### Base Memory

Base Memory is the private record of projects, people, preferences, and documents, stored and searched on owned hardware.

- **Memory you can inspect:** What AI knows about the user lives on Base Drive or Base Router as records they can open, correct, export, or delete.
- **The few facts, not your history:** When an assistant asks, Base searches locally and compiles the minimum context for that task. It never hands over the whole story.
- **Private details handled first:** Names, numbers, and sensitive details are swapped for placeholders on owned hardware, with a preview before anything leaves.
- **One memory, every assistant:** The same memory serves Claude, ChatGPT, and the local model. Users can switch assistants without starting over.

### Request flow

- The assistant asks Base, not the other way around.
- Base searches locally and returns the minimum.
- Requests outside sharing rules are declined.
- Every call is logged in plain language on Base.

### Works with

- **Claude:** add Base as a connector. Desktop talks over the user's own network; remote can bring memory to Claude on web and phone.
- **ChatGPT:** connect Base as a ChatGPT app. Workplaces can use a secure tunnel without opening a hole in the network.
- Anything else that speaks the same standard, including coding tools, agents, and note apps, can ask Base under the same rules.

### Tool set

- **search_memory:** Look something up in your memory. Returns permitted records with citations.
- **prepare_context:** Pack just enough context for one task. Compiles task-sized context within a hard budget.
- **get_source:** Show where a fact came from. Retrieves an approved excerpt from a cited source document.
- **preview_redaction:** Preview exactly what would be shared. Shows private details removed before anything sends.
- **propose_memory:** Suggest something new to remember. Nothing saves without confirmation on Base.

### Boundary

- A connector is not a firewall. It does not make ChatGPT or Claude local.
- Base Connect controls what its own tools return. The memory store stays home, and only tool results cross.
- Cleanup is careful, not magic. The connector returns the minimum, shows previews, and lets categories stay local-only.

## Use Cases

Route: `/use-cases`

### Where a local model earns its keep

The cloud is useful until it cannot be reached, or until the buyer would rather it never saw their data at all.

### Everyday help

The unglamorous daily wins. Short writing and clear explanations are what a small model is genuinely good at, and most days, that's most of the job.

- **Fix the email, keep the meaning** (Before you hit send): Paste a clumsy draft and get a cleaner one back. Tightening text is what a small model does best, and it never sees anyone's inbox but yours.
- **Get the short version** (The 30-page PDF): Summarize a long report on your own machine. Big documents take a little longer on small hardware, but the summary still beats reading page 19.
- **Explained like a neighbor would** (Official letters): A lease, an insurance letter, a school form, decoded into plain words. It explains; it doesn't advise.
- **Ask the same question five times** (Learning): A patient explainer that never judges and never gets tired. It can be wrong, so treat it as a tutor, not an authority.

### Private matters

Questions that are nobody's business. No account, no log, no server on the other end. Ask the things you'd never type into a cloud chatbot.

- **Health questions that stay yours** (The 3 a.m. question): Ask the awkward thing with nothing logged and no account attached. General information only. It is not a doctor.
- **Your budget, your business** (Money): Work through a budget or decode a confusing bill without a cloud service seeing your numbers.
- **Read the fine print, carefully** (Contracts): It can walk you through a clause in plain words. It is not legal advice; bring the big decisions to a professional.
- **A journal that talks back** (Reflection): Vent, reflect, plan. The transcript is a file on your disk that you can open, edit, or delete.

### No connection

The cloud is useful until you can't reach it. Planes, subways, trails, outages: a model that lives on your machine doesn't notice.

- **No signal, no problem** (On the subway): Underground with no bars and the cloud out of reach, a local model keeps working. Translate a sign, draft a reply, or summarize a long PDF on your commute, all on the device in your pocket.
- **Off-grid, miles from a tower** (In the mountains): No towers, no borrowed Wi-Fi. Ask first-aid questions, plan the next leg, or keep a trip journal on battery, fully offline. Help that never depends on a signal you don't have.
- **Remote, on battery, no data plan** (In the field): Deep in the woods for field work, log observations, draft notes, and ask general questions about what you're seeing. No data plan, no roaming, no round-trip to a server that can't hear you.
- **Cruising at 35,000 feet** (On a flight): Skip the pricey, patchy in-flight Wi-Fi. Outline a deck, tidy up your notes, or work through an email backlog with an assistant that runs entirely from your tray table.
- **No roaming, no problem** (Traveling abroad): Draft messages and translate short phrases without buying a data plan. Small-model translation is workable, not flawless.
- **Outages don't reach it** (When the internet is down): ISP outage, cloud outage, storm: the model on your laptop doesn't notice. Keep writing.

### Home & family

One machine for the whole house. No accounts to create, no profiles built on your kids, nothing leaving the living room.

- **Private by default, always on** (At home): Family documents, finances, and health questions never leave the house. A local model handles them on your own hardware, always on, with no subscription and nothing logged to someone else's cloud.
- **Help at the kitchen table** (Homework): Kids can ask questions without signing up for anything or being profiled. It can be wrong, so it is good for practice and explanations, not for copying answers.
- **One box, everyone on it** (The whole household): Base Router serves the whole house from one always-on box on your Wi-Fi. No per-person subscriptions.

### Work & school

Some documents aren't allowed to leave. Draft and summarize where the material already lives. For firms, clinics, and schools, there's a whole page.

- **Sensitive documents stay in** (At work): Contracts, patient notes, and source code never need to touch an outside API. Draft, review, and reason over the material where it already lives, inside your walls, under your control.
- **Meeting summaries without a vendor** (Internal notes): Summarize internal notes and drafts with no third party in the loop and nothing added to someone else's training data.
- **Student work stays in the building** (For teachers): Classroom help without sending student data to an outside service. A clean story for parents and the board.

### For tinkerers

You already know what localhost means. The part of the site where we can say GGUF. The models are files, the endpoint is local, and it's yours to mess with.

- **A local endpoint for your scripts** (It's just an API): Local models speak the same API shape your scripts already use. Point them at your own machine instead of a metered key.
- **Open WebUI, running locally** (Included UI): The drive ships Open WebUI wired to your own model, so you get a real chat interface from minute one.
- **Swap the model** (Your hardware, your rules): The models are ordinary GGUF files on disk. Replace them as the small-model field improves. No reinstall, no permission needed.
- **Apps against your own box** (Build on it): Our site-builder demo is the shape of it: one small model, one endpoint, your project on top.


## For Professionals

Route: `/professionals`

### You legally can't paste that into ChatGPT.

For firms, clinics, and schools, the confidential document is the whole job, and it is the one thing that cannot be handed to a public chatbot. Base Router is the upcoming hardware planned for that work.

- **Law firms (Client confidentiality):** Draft, summarize, and search privileged material on a box in the office. Nothing sensitive is pasted into a public chatbot.
- **Clinics (Patient trust):** Summarize notes and handle intake on hardware that stays in the building, with fewer outside vendors touching patient information.
- **Schools and districts (FERPA):** Student data stays in the building. The assistant supports classroom work without sending student work away.

### Built for accountability

- Activity records for every question and approval.
- Local-only mode for offices that need no outside connection.
- AI prepared for the field: intake, summarization, and search (in design).
- Team accounts with role-based access (in design).
- Optional secure remote access by user (in design).

### Sales model

Base Router is coming next and is not available to order. Its planned one-time price is $400, with no subscription required. A professional tier is in design with pilot customers.

### HIPAA honesty

HIPAA does not mandate data residency. The value of on-prem is narrower and concrete: it reduces business-associate agreements and sub-processor risk from sending protected data to a third party. Buyers should talk to counsel and use clear Offline Base documentation.

## Benchmarks

Route: `/benchmarks`

### Useful local AI starts with honest numbers.

The benchmark page uses Artificial Analysis screenshots dated May 30, 2026. No outside benchmark data is mixed in.

- Top score shown: 45.8.
- Visible range: 8.5-45.8.
- Output text volume: 3M-144M generated tokens during evaluation, context rather than a score.
- Small local models are around 15 on the same index as a hedged class position, not a measured product claim.

### Artificial Analysis Intelligence Index

A top-line score across the Artificial Analysis evaluation mix. Higher is better, and the spread explains Offline Base's honest limits. Source asset: `/benchmarks/intelligence-index.png`.

### Answer text produced during the index

Shows the amount of answer text produced during evaluation. It is context, not an intelligence score. Source asset: `/benchmarks/output-tokens.png`.

### Intelligence evaluations

Detailed panels for planning, coding, long documents, knowledge, scientific reasoning, instruction following, and image reasoning. Source asset: `/benchmarks/intelligence-evaluations.png`.

### Product takeaway

Bigger systems win the hardest tests, but local AI is still useful for private drafts, summaries, search, and focused work where ownership matters. The product promise is local by default, clear about outside help, and honest about limits before purchase.

## Local AI, Live

Route: `/try`

### Chat with a box, not a cloud.

A live chat with one model, Liquid LFM2.5 (1.2B), running on a real Jetson Orin Nano Super that Offline Base operates: the same board Base Router is built on. No AI company is in the loop, the chats are not stored, and replies stream at the pace the hardware really manages, about 38 tokens per second.

Conversations are capped at three messages because the demo runs on one small computer. The page exists to make a claim checkable: this is what AI on hardware you own feels like.

## Open Models

Route: `/open`

### Every model here is open.

A free chat, no account, with 10 open-weight models. Every model on the page has published weights that anyone can download and run on their own hardware.

These models run in the cloud, not on an Offline Base device. Requests go through OpenRouter with the `zdr: true` provider flag, so they are routed only to endpoints with a zero-data-retention policy; if a model loses that endpoint the request fails rather than falling back to a provider that logs. Every model on the page carries that flag. Zero retention is still a promise made by a third party, which is the distinction the page draws against running a model locally on hardware the user owns.

The page also frames these models as the cloud step in a local-first setup: hardware the user owns answers first, and an open cloud model picks up only the tasks that outgrow it.

Models available:

- **Nemotron 3.5 Lightning** (made by NVIDIA, medium at 30B, 1M context, handles very long documents). Good at: Fast, high-volume work. Wakes up just 3B of itself per word, so it answers fast for its size. This is what you get by default.
- **Kimi K3** (made by Moonshot, very large at 2.8T, 1M context, handles very long documents, reads images). Good at: Hard coding and deep reasoning. The largest model here at 2.8 trillion parameters. Reasons carefully, reads pictures, and is strong at hard coding problems.
- **DeepSeek V4 Flash** (made by DeepSeek, very large at 284B, 1M context, handles very long documents). Good at: Coding and agents. Only 13B of its 284B is active per word, so it stays quick despite its size. Built for coding and agent work.
- **Gemma 4 31B** (made by Google, medium at 31B, 262k context, handles very long documents, reads images). Good at: Writing, and reading images. Google's newest Gemma. Reads pictures as well as text, and can show its thinking.
- **Qwen3.6 27B** (made by Alibaba, medium at 27B, 262k context, handles very long documents, reads images). Good at: Reading images and video. Alibaba's dense mid-size model. Takes pictures and video as well as text, and can show its thinking.
- **Qwen3.8 27B** (made by Alibaba, medium at 27B, 262k context, handles very long documents, reads images). Good at: Coding and careful answers. The newest mid-size Qwen. Reads pictures and video, and thinks before it answers.
- **Muse Glimmer 30B** (made by Meta, medium at 30B, 131k context, handles long documents, reads images). Good at: Agents on modest hardware. From Meta Superintelligence Labs, shrunk down from a much bigger model so it can run on ordinary hardware. Reads pictures too.
- **Inkling Small** (made by Thinking Machines, very large at 276B, 524k context, handles very long documents, reads images). Good at: Quick multimodal answers. Thinking Machines Lab's smaller model. Wakes 12B of its 276B per word, so it answers quickly, and it reads pictures.
- **GLM 5.2** (made by Z.ai, large, 1M context, handles very long documents). Good at: Long, multi-step projects. A reasoning model built for long jobs: whole-project code work and tasks with many steps.
- **Qwen3.8 2.4T** (made by Alibaba, very large at 2.4T, 1M context, handles very long documents). Good at: The hardest questions. The open version of Qwen's largest model: 2.4 trillion parameters, of which 95B run for each word.

### What the chat can do

- **Reasoning:** models that think before answering stream that reasoning into a panel that opens while they work and collapses once the answer starts.
- **Tool calling:** models can call a lookup tool for Offline Base hardware facts, so answers about our products and local speeds come from the same data as the rest of this site rather than from guesswork. Every model in the catalogue supports tools.
- **Web search (model-decided):** the model can call OpenRouter's web-search tool when it judges a question needs the web, and the answer lists its sources. That search query goes to a search provider, which is outside the zero-retention guarantee, so the possibility is disclosed under the composer on every message.
- **File attachments (PDF and, on vision models, images):** PDFs are converted to text by OpenRouter's file-parser plugin, which is likewise outside the zero-retention guarantee and therefore opt-in. No model in this catalogue accepts files natively; Kimi K3, Gemma 4 31B, Qwen3.6 27B, Qwen3.8 27B, Muse Glimmer 30B and Inkling Small can read images.
- **File creation:** models can write a text file (notes, code, CSV) with a createFile tool; the download is assembled in the visitor's browser from the reply itself, so the file is never stored anywhere.
- **Thinking level:** a Quick / Balanced / Deep control in the composer sets how much reasoning the model is asked to spend. Balanced sends nothing and leaves each model at its own default.
- **Export:** the conversation can be downloaded as a Markdown file, assembled in the browser.

The chat is rate limited per IP address and capped per reply; it is a demo, not a general-purpose API.

## Build Your Own Drive

Route: `/build-your-drive`

### Design your Base Drive

Choose models and Offline Base sizes the drive, lists laptop requirements, and estimates price. Two models come loaded on the drive at launch; the rest are roadmap items.

Always-included planned safety model: Qwen3Guard (safety), 0.6B, 0.5 GB, 2 GB RAM.
App and offline documentation overhead: 3 GB.

### Model catalog

### LiquidAI LFM2.5 1.2B

- Status: On the drive at launch
- Parameters: 1.2B
- Disk estimate: 0.9 GB
- Minimum RAM: 4 GB
- GPU recommendation: CPU is fine
- Tier: tiny
- Summary: Fast, compact instruct model. Ships on every drive today.

### Gemma 4 E2B

- Status: On the drive at launch
- Parameters: ~2B
- Disk estimate: 3 GB
- Minimum RAM: 8 GB
- GPU recommendation: CPU is fine
- Tier: small
- Summary: Google's on-device model. The default everyday model, on every drive.

### MiniCPM5-1B

- Status: Roadmap
- Parameters: 1B
- Disk estimate: 0.8 GB
- Minimum RAM: 4 GB
- GPU recommendation: CPU is fine
- Tier: tiny
- Summary: Best intelligence-per-megabyte. Runs on weak laptops.

### Qwen 3.5 2B

- Status: Roadmap
- Parameters: 2B
- Disk estimate: 1.5 GB
- Minimum RAM: 4 GB
- GPU recommendation: CPU is fine
- Tier: small
- Summary: Snappy small general model, very long context.

### Gemma 4 E4B

- Status: Roadmap
- Parameters: ~4B
- Disk estimate: 4.5 GB
- Minimum RAM: 8 GB
- GPU recommendation: CPU is fine
- Tier: balanced
- Summary: The balanced pick. Has a reasoning mode for hard tasks.

### Qwen 3.5 9B

- Status: Roadmap
- Parameters: 9B
- Disk estimate: 6 GB
- Minimum RAM: 16 GB
- GPU recommendation: 8 GB VRAM
- Tier: large
- Summary: Strong reasoning. Wants a capable laptop or a GPU.

### Gemma 4 12B

- Status: Roadmap
- Parameters: 12B
- Disk estimate: 8.1 GB
- Minimum RAM: 16 GB
- GPU recommendation: 10 GB VRAM
- Tier: large
- Summary: More capable, for 16GB+ machines.

### Qwen 3.6 27B

- Status: Roadmap
- Parameters: 27B
- Disk estimate: 16 GB
- Minimum RAM: 32 GB
- GPU recommendation: 16 GB VRAM
- Tier: xlarge
- Summary: Top-tier. Really a model for a desktop with a serious GPU.

### Example estimates

- Everyday default (Gemma 4 E2B): 6.5 GB on drive, 32 GB USB recommended, 8 GB RAM, estimated $35.
- Heavier build (Gemma 4 E2B + Qwen 3.5 9B + Gemma 4 12B): 20.6 GB on drive, 32 GB USB recommended, 16 GB RAM, 10 GB GPU recommended, estimated $55.

Estimates use Q4 model footprints and mock build pricing, not checkout totals.

## Shop And Cart

Route: `/store`

### Everything in the store is coming soon.

No product is orderable yet. Base Drive will be the first to launch, at $59. Base Drive Lite and Pro are previews, and Base Router follows. When ordering opens, checkout will be handled securely by Shopify.

If Shopify environment variables are missing, the store page says the store is not connected yet and asks for `SHOPIFY_STORE_DOMAIN` and `SHOPIFY_STOREFRONT_ACCESS_TOKEN` in `.env.local`.

The cart page shows product, price, quantity, total, subtotal, and a Shopify-hosted checkout link. Empty carts point users back to the store.

## Links

Route: `/links`

### Offline Base links

Private AI that runs on your own laptop. Base Drive and Base Router are coming soon.

- Base Drive, $59, coming soon - Private AI on a USB stick.
- Build your own drive - pick models, get size, specs, and price.
- Try the AI live for free - chat with a model running on our box.
- What's on the drive - models, chat app, and readable memory.
- How it works - plug in, run one file, chatting in minutes.
- Base Router, coming next - preview the upcoming always-on product.
- Verified profiles: X, Instagram, YouTube.

## Source Documents

The following Markdown files are included in full and nested under this section so agents can traverse the research without opening separate files.

## Source Document: Two Product Routes

Source file: `docs/TWO_PRODUCT_ROUTES.md`

<!-- source-document-start -->

### Offline Base: Two Separate Product Routes

Decision date: June 13, 2026

#### Decision

Build both ideas, but present and operate them as separate products:

1. **Base Memory:** a private memory, retrieval, and PII gateway for Claude,
   ChatGPT, and other assistants.
2. **Base Lab:** a local AI workstation for running and experimenting with
   models on Jetson hardware.

They can share hardware engineering, installation, storage, authentication,
and update infrastructure. They should not share the same default user
experience or security promise.

| | Route 1: Base Memory | Route 2: Base Lab |
|---|---|---|
| Primary buyer | Mainstream and professional users | Developers, enthusiasts, researchers |
| Main outcome | Better cloud AI context with less private data exposed | Run and test AI locally on owned hardware |
| Default interface | Simple Base app plus Claude and ChatGPT connectors | Model dashboard, local chat, API, and developer tools |
| Core value | Memory, retrieval, minimization, redaction, portability | Control, experimentation, offline inference, no per-token fee |
| Default security posture | Locked-down appliance | User-controlled development environment |
| Cloud requirement | Optional, when using Claude or ChatGPT | None for qualified local models |
| Support promise | Guided and predictable | Flexible, with compatibility boundaries |

#### Route 1: Base Memory

##### Product promise

> Your Base remembers what matters, finds only the details needed for the
> current task, and removes private information before optional outside help.

Base Memory is not primarily a local chatbot. It is a private context and
privacy layer that sits between a person's information and the AI they choose
to use.

##### Target users

- Knowledge workers with confidential documents
- Students with notes, readings, and course history
- Families managing household records
- Older adults who need help understanding letters and forms
- Small offices that want shared memory without uploading an entire drive

##### Core capabilities

- Encrypted local memory
- Structured records with human-readable Excel export
- Optional selected-row Google Sheets import or mirror
- Full-text and semantic retrieval
- Context budgeting and deduplication
- Local PII detection and pseudonymization
- Source citations
- Consent preview before outside processing
- Portable Memory Pack for moving between Base products

##### Three ways to use it

###### 1. Base Gateway

The user chats through the Base app. Base Memory:

1. Receives the prompt.
2. Searches local memory.
3. Compiles the minimum context.
4. Detects and removes configured PII.
5. Shows an outbound preview when required.
6. Calls Claude, OpenAI, or another approved provider API.
7. Restores local placeholders where safe.
8. Returns the answer with memory citations and a privacy receipt.

This is the strongest privacy path because the Jetson processes the request
before the provider receives it.

It requires provider API access. A ChatGPT or Claude subscription should not be
described as including API usage; API authentication and billing are separate.

###### 2. Native Claude connector

Expose Base Memory as an MCP server.

For a local-first desktop setup, package it as a Claude Desktop extension that
connects to the Base over the local network. Claude Desktop supports local MCP
servers and packaged desktop extensions.

For Claude web, mobile, Cowork, and account-wide availability, offer a remote
MCP connector. Anthropic documents remote custom connectors across Free, Pro,
Max, Team, and Enterprise plans, with one connector for Free users. Remote
connector requests originate from Anthropic's cloud, so the MCP endpoint must
be reachable from Anthropic's infrastructure.

Use the connector for:

- `search_memory`
- `prepare_context`
- `get_source`
- `preview_redaction`
- `propose_memory`

Start with read-only tools. Saving or deleting memory should require explicit
confirmation in the Base interface.

###### 3. Native ChatGPT app

Expose the same MCP tools to ChatGPT through an Apps SDK app or custom MCP app.

OpenAI currently documents:

- ChatGPT apps backed by MCP servers
- Custom MCP apps for Business and Enterprise/Edu workspaces
- OAuth with per-tool scopes
- Secure MCP Tunnel for reaching private or on-premises MCP servers without
  opening a public inbound endpoint

The Secure MCP Tunnel is the preferred Jetson deployment when the customer's
ChatGPT workspace supports it. The tunnel client runs beside Base Memory,
makes an outbound HTTPS connection, and leaves the private MCP address inside
the customer's network.

For broader future distribution, an approved ChatGPT app-directory listing can
be evaluated separately. Do not make public-directory approval a dependency
for the first product.

##### Critical connector limitation

A connector is not a guaranteed firewall.

ChatGPT and Claude decide when a tool is relevant based on the prompt, tool
description, and system instructions. A user can also type private information
into the provider chat before Base Memory is called. At that point the original
text has already reached the provider.

Therefore:

| User need | Correct interface |
|---|---|
| "Let Claude or ChatGPT search my minimized local memory" | Native MCP connector |
| "Make sure every outgoing prompt is redacted first" | Base Gateway |
| "Keep the whole task off the cloud" | Base Lab local model |

Marketing must preserve this distinction. The connector keeps the memory store
local and controls what its tools return. It does not make Claude or ChatGPT
local.

##### MCP tool surface

Keep the public connector small:

| Tool | Purpose | Initial mode |
|---|---|---|
| `search_memory` | Return a short list of permitted records and citations | Read |
| `prepare_context` | Compile a token-budgeted context package for a task | Read |
| `get_source` | Retrieve an approved excerpt from a cited source | Read |
| `preview_redaction` | Show how a proposed payload would be minimized and pseudonymized | Read |
| `propose_memory` | Submit a candidate memory for review in the Base app | Proposal only |
| `export_memory_pack` | Request an export that the user approves locally | Local confirmation |

Do not expose:

- Raw database queries
- "Return all memory"
- Unscoped file access
- Direct deletion without local confirmation
- The pseudonym replacement vault
- Administrative credentials

##### Route 1 MVP

1. Build the encrypted memory service and stable record schema.
2. Implement deterministic retrieval and context compilation.
3. Implement PII preview and pseudonymization.
4. Ship the Base Gateway with one provider API.
5. Expose the read-only MCP tool set.
6. Package a Claude Desktop local connector.
7. Test ChatGPT through Secure MCP Tunnel in an eligible workspace.
8. Add remote Claude and ChatGPT deployment only after OAuth, permissions, and
   audit logging are qualified.

##### Route 1 business model

Possible packaging:

- Hardware purchase
- Base Memory software included for one user
- Paid family or team accounts
- Optional annual updates and support
- No required cloud subscription from Offline Base
- Provider API costs paid directly by the user or organization

The strongest differentiator is not cheaper inference. It is owned memory that
works across providers.

#### Route 2: Base Lab

##### Product promise

> A small local AI workstation you own, with the tools to install, compare, and
> use models without sending every task to the cloud.

Base Lab serves buyers who see the Jetson as a general AI computer rather than
only a privacy appliance.

##### Target users

- Developers and technical students
- Local AI enthusiasts
- Researchers and educators
- Makers working with cameras, robotics, speech, or sensors
- Small teams testing local inference before a larger deployment

##### Core capabilities

- Curated model catalog
- One-click install, update, start, stop, and remove
- Ollama or CUDA `llama.cpp` runtime
- Local web chat
- Authenticated OpenAI-compatible LAN endpoint
- Model compatibility and memory estimates
- Tokens-per-second, time-to-first-token, memory, and temperature measurements
- Prompt and model profiles
- Document, speech, OCR, and vision starter workflows
- Optional coding clients such as Aider, Pi, Cline, or OpenCode from a laptop
- Backup and restore for models and configuration

##### Hardware reality

The current Jetson Orin Nano Super has 8 GB of shared memory. The existing
research recommends treating 2B to 4B quantized models as the comfortable
range. Some 7B models may run, but model weights, context cache, operating
system, runtime, and other services compete for the same memory.

Base Lab should sell a qualified experience, not "run any model."

Recommended tiers:

| Product | Hardware | Position |
|---|---|---|
| Base Lab Mini | Raspberry Pi 5 | Learning, very small models, gateway, and utilities |
| Base Lab | Jetson Orin Nano Super 8 GB | Small local language, vision, speech, and coding models |
| Base Lab Pro | Jetson Orin NX 16 GB or AGX Orin | Larger models, longer context, and heavier agents |

##### Experience

The first screen should answer:

- What can this hardware run?
- What is installed?
- What is running now?
- How much memory is free?
- Which apps can connect?
- Is this model local or outside?

The user chooses a workload, not only a model:

- Private chat
- Coding assistant
- Document search
- Transcription
- Translation
- OCR and document parsing
- Image detection and segmentation
- Lightweight robotics

##### Route 2 MVP

1. Pin one JetPack and operating-system image.
2. Qualify one runtime and three small models.
3. Build install, start, stop, remove, and health controls.
4. Add local chat and authenticated LAN API.
5. Record hardware measurements for every supported configuration.
6. Add a safe model import path with disk and memory checks.
7. Publish a dated compatibility catalog.
8. Add specialist speech, OCR, vision, and coding packs.

##### Route 2 business model

Possible packaging:

- Hardware margin
- Free local model manager
- Paid curated model and workflow packs
- Pro hardware tiers
- Optional support plan
- Paid training or team deployment services

Avoid a required subscription for core local operation. Ownership is the
product's main advantage.

#### Why the routes must remain separate

##### Different trust promises

Base Memory promises controlled access to sensitive information. Base Lab
promises freedom to install and modify software. Unrestricted developer access
weakens the assurance that Base Memory's privacy controls have not been
changed.

##### Different support burden

Base Memory should have a small, tested configuration. Base Lab users will
install unsupported models, alter services, exhaust storage, and expect
technical controls.

##### Different interfaces

Base Memory should hide model names and begin with tasks such as "Search my
records" or "Ask with private details removed." Base Lab should expose models,
runtimes, ports, memory, logs, and benchmarks.

##### Different failure handling

Base Memory should fail closed when identity, permission, or redaction checks
fail. Base Lab can report a model error and let a technical user debug it.

#### Shared foundation

Build these components once:

- Signed base operating-system image
- Device identity and local authentication
- Encrypted storage and backup
- Update and rollback service
- Local web dashboard shell
- Network discovery
- Runtime health and resource monitoring
- Audit event format
- Memory Pack and configuration export
- Authenticated local API gateway

Keep these separate:

| Base Memory | Base Lab |
|---|---|
| Memory database and PII vault | User model files and experiments |
| Locked service account | Developer workspace |
| Stable release channel | Faster lab release channel |
| Narrow MCP tools | General local API |
| Privacy audit log | Runtime and benchmark log |
| Fail-closed policy | Debuggable errors |

#### Packaging recommendation

Use one hardware platform with two software editions:

- **Base Memory Edition**
- **Base Lab Edition**

Do not ship a casual toggle that converts a locked privacy appliance into an
unrestricted development machine. Offer a deliberate reimage process:

1. Export an encrypted Memory Pack.
2. Confirm that changing editions alters the security posture.
3. Erase the system partition.
4. Install the other signed image.
5. Restore only compatible user data.

A future higher-memory device could run both in isolated virtual machines or
containers, but the 8 GB Orin Nano does not have enough headroom to make that
the first product.

#### Brand structure

Recommended:

```text
Offline Base
├── Base Memory
│   ├── Base Gateway
│   ├── Claude connector
│   └── ChatGPT app
└── Base Lab
    ├── Local model manager
    ├── Local chat
    └── Local API
```

Alternative hardware naming:

- Base Router, Memory Edition
- Base Router, Lab Edition

Avoid calling both products simply "Base Router." The buyer should know whether
they are purchasing a privacy appliance or an AI development computer.

#### Recommended order

##### First: Base Memory

Why:

- Stronger differentiation
- Clear professional privacy problem
- Works with Claude and ChatGPT instead of competing directly with them
- Creates provider-independent user memory
- The PII and minimum-context story is easier to explain and defend

##### In parallel, narrowly: Base Lab foundation

Build the runtime manager, local chat, resource monitor, and three-model
compatibility catalog. Do not expand into a large app store until the core
memory product is stable.

##### Then: connect the two without combining the promises

Base Memory may use a qualified local Base Lab model as one possible answer
provider. This keeps a task fully local when the small model is capable.

The dependency should be one-directional:

```text
Base Memory -> qualified local inference API

Base Lab -X-> unrestricted access to the memory vault
```

Base Lab never receives automatic access to private memory. The user must
create an explicit, scoped connection.

#### Sources

OpenAI:

- [Build an MCP server for ChatGPT Apps](https://developers.openai.com/apps-sdk/build/mcp-server)
- [MCP and Connectors](https://developers.openai.com/api/docs/guides/tools-connectors-mcp)
- [Secure MCP Tunnel](https://developers.openai.com/api/docs/guides/secure-mcp-tunnels)
- [Apps SDK authentication](https://developers.openai.com/apps-sdk/build/auth)
- [Developer mode and MCP apps in ChatGPT](https://help.openai.com/en/articles/12584461-developer-mode-and-mcp-apps-in-chatgpt)
- [ChatGPT and API billing are separate](https://help.openai.com/en/articles/8156019-how-can-i-move-my-chatgpt-subscription-to-the-api)

Anthropic:

- [Custom connectors using remote MCP](https://support.claude.com/en/articles/11175166-get-started-with-custom-connectors-using-remote-mcp)
- [Local MCP servers and Claude Desktop extensions](https://support.claude.com/en/articles/10949351-getting-started-with-local-mcp-servers-on-claude-desktop)
- [Claude API MCP connector](https://platform.claude.com/docs/en/agents-and-tools/mcp-connector)
- [Claude subscriptions and API billing are separate](https://support.claude.com/en/articles/9876003-i-have-a-paid-claude-subscription-pro-max-team-or-enterprise-plans-why-do-i-have-to-pay-separately-to-use-the-claude-api-and-console)

<!-- source-document-end -->

## Source Document: Private Memory Router Architecture

Source file: `docs/PRIVATE_MEMORY_ROUTER_ARCHITECTURE.md`

<!-- source-document-start -->

### Offline Base: Private Memory Router Architecture

Decision date: June 13, 2026

For the separation between the memory product and the local-model product, see
[Two Separate Product Routes](./TWO_PRODUCT_ROUTES.md).

#### Recommendation

The idea is strong and gives Base Router a concrete reason to exist beyond
hosting a small chatbot.

The product should be:

> A private memory router that finds the few facts a task needs, removes
> sensitive details when required, and sends only the minimum useful context to
> the model doing the work.

This can reduce prompt size, improve response speed, and make outside model use
safer. Three design corrections are important:

1. Use a small number of narrow workers, not an autonomous swarm of agents.
2. Make a spreadsheet the user-readable view and interchange format, not the
   only database.
3. Make retrieval and PII handling explicit local services. Do not depend on a
   general language model to enforce the privacy boundary.

#### What actually creates the savings

##### Token savings

The main token reduction comes from selective retrieval and context
compilation:

- Search structured memory before calling the main model.
- Return atomic records instead of entire conversations or documents.
- Deduplicate overlapping records.
- Include only the fields needed for the current task.
- Set a hard context budget.
- Summarize older material once, then retrieve the summary when appropriate.

Sub-agents do not automatically save tokens. They increase total usage when
each worker receives the same conversation, repeats the same search, or writes
long handoff messages.

##### Latency improvements

Latency improves when:

- Identity, permission, date, category, and sensitivity filters run before
  semantic search.
- Full-text and vector search run in parallel.
- Only one model generation is required for a normal request.
- Memory curation happens after the user receives the answer.
- The Jetson keeps one qualified model loaded and serializes generation on
  8 GB hardware.

Latency gets worse when several agents call the same model one after another.
The normal path should be one orchestrator decision, deterministic retrieval,
and one answer generation.

##### Privacy improvements

Privacy improves when:

- The authoritative memory remains encrypted on the laptop or Base Router.
- Every record has an owner, workspace, sensitivity level, and sharing scope.
- The system retrieves the minimum required information.
- PII detection runs locally.
- Outside requests are pseudonymized or redacted before the consent screen.
- The user sees the exact sanitized payload before approving it.
- The identity mapping never leaves the local device.

#### Recommended system shape

```text
User request
     |
     v
Orchestrator
  - identifies task
  - selects tools
  - sets context budget
     |
     +-------------------------+
     |                         |
     v                         v
Memory retrieval          Task-specific tools
  - permission filter       - OCR
  - metadata filter         - transcription
  - full-text search        - translation
  - vector search           - document parsing
  - reranking
     |                         |
     +------------+------------+
                  |
                  v
Context compiler
  - deduplicates
  - selects fields
  - applies token budget
  - attaches provenance
                  |
                  v
Privacy gate
  - local task: least-privilege context
  - outside task: detect and replace PII
                  |
                  v
Local model or approved outside model
                  |
                  v
Answer with sources and privacy status
                  |
                  v
Asynchronous memory curator
  - proposes new records
  - merges duplicates
  - asks before saving sensitive facts
```

#### Worker boundaries

The implementation can use agent-style workers, but each one should have a
narrow contract.

##### 1. Orchestrator

Responsibilities:

- Understand the user's requested outcome.
- Choose the minimum required tools.
- Decide whether memory is needed.
- Set a maximum number of records and context tokens.
- Decide whether the task can remain local.

The orchestrator should not read the entire memory store. It receives only
search metadata and selected results.

##### 2. Memory retriever

Responsibilities:

- Enforce user and workspace permissions before search.
- Search structured fields, full text, and optional embeddings.
- Apply date, category, source, and sensitivity filters.
- Return short records with provenance and relevance scores.

This should primarily be a deterministic service, not an LLM agent. A small
model may translate an ambiguous user request into filters, but it should not
bypass the service's permission rules.

##### 3. Privacy guard

Responsibilities:

- Detect configured PII types with rules and local entity models.
- Apply allowlists and organization-specific recognizers.
- Replace sensitive values with stable placeholders.
- Keep the replacement map in an encrypted local vault.
- Produce a reviewable sanitized payload.

Microsoft Presidio is a practical starting framework because it supports
rule-based and model-based recognizers, custom entity types, and local
operation. It still requires evaluation because PII detectors produce both
false positives and false negatives.

##### 4. Task worker

Responsibilities:

- Perform the actual writing, comparison, explanation, planning, or other
  requested work.
- Use only the compiled context.
- Cite the memory records or source documents used.
- Return uncertainty when the retrieved memory does not support an answer.

This can be the same loaded general model used by the orchestrator. Separate
worker identities do not require separate model weights.

##### 5. Memory curator

Responsibilities:

- Inspect the completed interaction after the answer is returned.
- Propose small, durable facts rather than saving the whole conversation.
- Merge duplicates and supersede outdated facts.
- Assign source, confidence, sensitivity, and expiration.
- Ask before saving sensitive or surprising information.

The curator should be asynchronous and lower priority so it does not delay the
answer.

#### Memory storage decision

##### Canonical store: encrypted SQLite

Use an encrypted SQLite database as the authoritative memory store.

Why:

- Atomic updates are safer than rewriting an Excel workbook.
- SQLite FTS5 provides efficient local full-text search.
- Structured filters can run before semantic retrieval.
- Stable IDs, relationships, access rules, deletion records, and schema
  migrations are straightforward.
- The same database design runs on Linux, macOS, and Windows.
- The database can be backed up as a consistent snapshot.

SQLCipher is one option for transparent encrypted SQLite storage, subject to
license and packaging review. The encryption key should come from the operating
system keychain, secure hardware where available, or a user recovery secret. It
should not be stored beside the database.

##### User-readable view: Excel workbook

Generate `Memory.xlsx` from the database and allow reviewed edits to be
imported.

The workbook is valuable because it is:

- Familiar
- Inspectable
- Editable
- Easy to archive
- Portable between Base Drive, Base Router, and another laptop

It should not be the live database because workbook edits do not provide strong
transactions, row-level authorization, reliable concurrent writes, or an
efficient search index.

##### Optional cloud view: Google Sheets

Google Sheets should be an explicit opt-in integration, not the default memory
location.

Google's Sheets API supports multiple sheets containing rows, columns, and
cells. Google Drive can export a spreadsheet as `.xlsx`, which makes it useful
as an interchange option. It also means the selected memory is stored and
synced through Google rather than remaining entirely local.

Recommended modes:

| Mode | Authoritative memory | Privacy position |
|---|---|---|
| Local-only | Encrypted database on laptop or Base Router | Strongest |
| Portable | Local database plus encrypted Memory Pack export | Strong |
| Google mirror | Local database plus selected synchronized rows | User-approved cloud copy |
| Google import | Read a chosen Sheet, import selected rows, then disconnect | Temporary outside source |

Do not silently synchronize all memory to Google Sheets. The integration should
show which tabs and fields will be copied, require a separate approval, and be
disabled in local-only mode.

#### Workbook design

Tabs should improve human navigation without becoming separate incompatible
databases.

Recommended tabs:

| Tab | Purpose |
|---|---|
| `Profile` | User-approved stable information and communication preferences |
| `People` | Contacts and relationships |
| `Projects` | Work projects, classes, cases, or ongoing goals |
| `Preferences` | Reusable choices, styles, routines, and defaults |
| `Records` | Atomic facts, decisions, events, and observations |
| `Sources` | Documents, messages, recordings, and source hashes |
| `Sharing` | Sensitivity, consent, retention, and outside-use rules |
| `Change Log` | Imports, edits, merges, deletions, and device transfers |

The `Records` tab should remain the shared core. Category-specific tabs can be
generated filtered views. Otherwise, adding a new category forces every search
and migration tool to understand another layout.

#### Record schema

Each memory record should represent one durable claim or event.

| Field | Purpose |
|---|---|
| `memory_id` | Stable UUID used across every device and export |
| `owner_id` | Person who owns the memory |
| `workspace_id` | Home, class, project, client, or team boundary |
| `category` | Profile, person, preference, project, event, decision, or other controlled value |
| `subject` | Person, project, class, account, or topic the record concerns |
| `summary` | Short retrieval-friendly statement |
| `details` | Optional structured JSON or longer explanation |
| `source_id` | Link to the supporting source |
| `source_excerpt` | Minimal supporting text where appropriate |
| `source_hash` | Detects changed source material |
| `confidence` | Confirmed, inferred, or uncertain |
| `sensitivity` | Public, internal, private, restricted, or secret |
| `pii_types` | Detected entity classes |
| `consent_scope` | Local only, named devices, named people, or approved outside use |
| `valid_from` | When the fact became true |
| `valid_until` | When the fact stopped being true |
| `created_at` | Original creation time |
| `updated_at` | Latest reviewed update |
| `device_origin` | Product that created the record |
| `status` | Active, superseded, disputed, expired, or deleted |

Do not use spreadsheet row number as identity. Rows move. Every record needs a
stable `memory_id`.

#### Retrieval pipeline

##### Step 1: Establish identity and scope

Before search, resolve:

- Active user
- Allowed workspaces
- Device
- Requested task
- Whether outside processing is permitted

This prevents cross-user memory leakage before relevance scoring begins.

##### Step 2: Apply cheap filters

Filter by:

- Owner and workspace
- Record status
- Category
- Date range
- Source type
- Sensitivity and consent scope

##### Step 3: Run hybrid search

Run in parallel:

- SQLite FTS5 keyword and phrase search
- Optional local embedding search
- Exact lookup for names, IDs, dates, and known fields

Merge the results and rerank them. Embeddings should be treated as a
regenerable index, not the portable source of truth.

##### Step 4: Compile minimal context

The context compiler should:

- Select a small top set.
- Remove duplicates and superseded facts.
- Prefer confirmed records with sources.
- Include only needed columns.
- Preserve stable record IDs for citations.
- Stop at a configured token budget.

If the evidence is weak, the system should ask a question instead of widening
the search indefinitely.

#### PII and minimum-disclosure pipeline

The Jetson should distinguish three operations.

##### 1. Classification

Label records and retrieved fields by sensitivity and PII type. Classification
does not modify the authoritative local record.

##### 2. Context minimization

Remove irrelevant fields even for local inference. For example, a scheduling
question may need a person's name and availability but not their address,
account number, or unrelated history.

This is what saves the most tokens while preserving usefulness.

##### 3. Redaction or pseudonymization

Before an outside request:

- Replace names and identifiers with stable tokens such as `[PERSON_01]`.
- Generalize unnecessary dates or locations.
- Remove hidden metadata and unused source excerpts.
- Keep the local replacement map encrypted.
- Show the user what will leave.
- Reinsert local names only after the outside response returns, when safe.

Redaction is not sufficient by itself. A combination of facts can identify a
person even when names are removed. The privacy guard must also support
minimum disclosure and user review.

#### Portable Memory Pack

The transfer unit between Base products should be an encrypted bundle:

```text
Base-Memory-2026-06-13/
  manifest.json
  memory.db
  Memory.xlsx
  schema.json
  sources/
  checksums.sha256
  README.txt
```

The manifest should include:

- Schema version
- Exporting product and software version
- User and workspace IDs
- Creation time
- Included categories and sources
- Encryption and key-recovery method
- Required migrations

Import behavior:

1. Verify signature and checksums.
2. Unlock with the user's recovery secret or approved device.
3. Validate the schema version.
4. Merge by stable UUID.
5. Preserve deletions and superseded records.
6. Resolve conflicts visibly.
7. Rebuild FTS and embedding indexes locally.
8. Produce an import report.

This makes the user's memory portable without tying it to a particular model,
agent framework, or Base product.

#### Product behavior

##### Base Drive

- The laptop hosts the encrypted memory database.
- `Memory.xlsx` provides the readable and editable view.
- Base Drive carries installers, schema, and an encrypted Memory Pack backup.
- Retrieval and PII processing run on the laptop.
- A move to Base Router uses the same export and import format.

##### Base Router

- The Jetson or attached NVMe hosts the authoritative database.
- Phones and laptops access memory through a permissioned local service.
- The Jetson performs retrieval, context compilation, and PII processing.
- Multiple users receive separate identities and workspaces.
- Memory exports are explicit and logged.

##### Google Sheets integration

- The user chooses the Sheet and tabs.
- The Base previews imported or exported fields.
- Synchronization uses stable `memory_id` values.
- Only approved rows and columns synchronize.
- Local-only mode disables the connector.
- API errors, quota errors, and conflicts cannot block access to local memory.

#### What not to build

Avoid:

- Several autonomous agents continuously talking to one another
- One agent with unrestricted access to every file and credential
- Every worker receiving the full chat history and full spreadsheet
- Excel or Google Sheets as the only live database
- Automatic cloud synchronization under a "backup" label
- Saving every message as permanent memory
- Treating embeddings as facts
- Redacting only names while leaving identifying combinations intact
- Letting retrieved memory contain executable instructions
- Running several large models concurrently on an 8 GB Jetson

Memory content must be treated as untrusted data. A retrieved row that says
"ignore prior instructions" is a record, not an instruction to the agent.

#### MVP sequence

##### Phase 1: Local memory foundation

- Encrypted SQLite store
- Stable schema and UUIDs
- Manual add, edit, approve, delete, export, and import
- FTS5 search
- Generated Excel workbook
- One orchestrator and one model generation per request

##### Phase 2: Privacy gateway

- Presidio with selected local recognizers
- Sensitivity and consent scopes
- Context minimization
- Pseudonymization vault
- Review screen before outside requests
- Local approval record

##### Phase 3: Better retrieval

- Local embedding index
- Hybrid ranking
- Source citations
- Token-budgeted context compiler
- Asynchronous memory curator

##### Phase 4: Portability and optional sync

- Signed encrypted Memory Pack
- Base Drive to Base Router migration
- Conflict handling and deletion propagation
- Opt-in Google Sheets import and selected-row mirror

##### Phase 5: Persona-specific views

- Student course memory
- Knowledge-worker project memory
- Older-adult everyday records and trusted-helper setup

All three should use the same memory service and schema.

#### Qualification metrics

Measure the complete system, not only model tokens per second.

##### Retrieval

- Recall of required facts in a representative task set
- Unsupported-memory rate
- Duplicate and stale-record rate
- Cross-user and cross-workspace leakage count

##### Efficiency

- Context tokens before and after retrieval
- Number of model calls per request
- Time to first useful answer
- Search and PII processing latency
- Peak Jetson memory

##### Privacy

- PII false-negative rate by entity class
- PII false-positive rate
- Percentage of outside requests reviewed
- Amount of original text removed before outside processing
- Recovery and deletion correctness across devices

##### Reliability

- Import and export round-trip fidelity
- Recovery after power loss during a write
- Google synchronization conflict rate
- Correct behavior with no internet
- Correct behavior when a Sheet or Google account becomes unavailable

#### Product positioning

Do not lead with spreadsheets, databases, embeddings, or sub-agents.

Lead with:

> Your Base remembers what matters, finds only what the current task needs, and
> removes private details before outside help is allowed.

Supporting messages:

- **Memory you can inspect:** Open it, correct it, move it, or delete it.
- **Less context, better focus:** The model receives the relevant facts, not
  your entire history.
- **Portable by design:** Move your memory from Base Drive to Base Router
  without starting over.
- **Private before powerful:** Sensitive details are handled locally before any
  optional outside request.

#### Source notes

- [SQLite FTS5](https://sqlite.org/fts5.html) provides local full-text search
  over SQLite virtual tables.
- [SQLite Online Backup API](https://sqlite.org/backup.html) provides
  consistent database snapshots while an application remains active.
- [SQLCipher](https://www.zetetic.net/sqlcipher/) provides encrypted SQLite
  database storage and requires product-level license review.
- [Microsoft Presidio Analyzer](https://microsoft.github.io/presidio/analyzer/)
  combines rule-based and model-based PII recognizers and supports custom
  recognizers.
- [Presidio recognizer guidance](https://microsoft.github.io/presidio/analyzer/developing_recognizers/)
  explicitly notes the need to evaluate false positives and false negatives.
- [Google Sheets API concepts](https://developers.google.com/workspace/sheets/api/guides/concepts)
  define spreadsheets as collections of sheets containing rows, columns, and
  cells.
- [Google Drive export formats](https://developers.google.com/workspace/drive/api/guides/ref-export-formats)
  support exporting Google Sheets to Microsoft Excel `.xlsx`.
- [Google Sheets API limits](https://developers.google.com/workspace/sheets/api/limits)
  document per-minute API quotas, so cloud synchronization should not sit on
  the critical local request path.
- [Google Drive for desktop](https://support.google.com/drive/answer/10838124)
  synchronizes files between a computer and Google Drive, which changes the
  local-only privacy boundary.

<!-- source-document-end -->

## Source Document: AI Model Use Case Positioning

Source file: `docs/AI_MODEL_USE_CASE_POSITIONING.md`

<!-- source-document-start -->

### Offline Base: Specialist AI Model and Use-Case Positioning

Research date: June 13, 2026

#### Executive decision

Offline Base should not be positioned as a box containing a long list of AI
models. Model names change too quickly, and most buyers do not care which
checkpoint performs a task.

The stronger position is:

> Offline Base is a private model router. It chooses the smallest local tool
> that can do the job, keeps sensitive work on hardware you control, and asks
> before outside help is used.

This turns the product's current privacy and consent story into a practical
system:

1. A local model identifies the task.
2. A specialist handles redaction, OCR, search, speech, translation, or vision.
3. A general model explains or combines the result.
4. If the task needs a larger outside model, sensitive information can be
   removed locally before the consent prompt appears.

The product should sell outcomes and workflow packs, not model inventory.

#### Product roles

The repo currently uses two naming systems for the same box products:

| Public navigation | Product data and pages | Hardware |
|---|---|---|
| Base Router Lite | Base Mini | Raspberry Pi 5, 8 GB |
| Base Router | Base One | Jetson Orin Nano Super, 8 GB |

This should be normalized before launching model-specific use cases. The
recommendation below uses **Base Router Lite** and **Base Router** because those
names support the model-routing position.

##### Base Drive Lite

**Role:** A portable offline utility kit for older or lower-memory laptops.

Best-fit capabilities:

- Personal information redaction
- Basic OCR
- Voice activity detection and short transcription
- Local file search
- Small text translation models
- Prompt and content safety checks

Do not position it as a replacement for a large general chatbot.

##### Base Drive

**Role:** A personal offline AI toolkit that is matched to the buyer's laptop.

Best-fit capabilities:

- Document intake, search, summarization, and redaction
- Meeting and interview transcription
- Text and speech translation
- Image understanding on capable laptops
- Coding and writing assistance
- Optional specialist packs selected during setup

The laptop provides the compute, so this product can support a wider range of
models than either fixed 8 GB box.

##### Base Drive Pro

**Role:** A portable local AI workbench for developers, researchers, and
technical field teams.

Best-fit capabilities:

- Local model server
- Custom model and adapter installation
- GPU-accelerated segmentation and detection
- Robotics experimentation
- Batch document processing
- Workflow development before deploying to a Base Router

##### Base Router Lite

**Role:** An always-on privacy and knowledge appliance for a home or small
office.

Best-fit capabilities:

- Local redaction before cloud escalation
- Shared document OCR and search
- Small-model classification and extraction
- Audio transcription queues
- Translation of text and documents
- Safety, prompt-injection, and policy checks

The Raspberry Pi should run short, focused jobs. It should not be sold for
real-time video analysis, large multimodal models, or general robotics control.

##### Base Router

**Role:** An always-on local model router for a family, office, workshop, or
field site.

The Jetson Orin Nano Super has 8 GB of memory, up to 67 INT8 TOPS, and a 7 to
25 watt power range. That makes it well suited to small language models,
specialist vision models, OCR, speech, and lightweight robotics. Its memory is
still the hard limit.

Best-fit capabilities:

- Everything in Base Router Lite
- GPU-accelerated OCR and document vision
- Object detection, segmentation, and image classification
- Shared speech and translation services
- Small vision-language models
- Lightweight robot perception and action models

Position it as a router that loads the right specialist for a task, not as a
server that keeps every model in memory at once.

##### Future Base Edge or Base Robotics

A higher-memory product is needed before Offline Base can credibly support the
latest general-purpose robotics models, several live camera streams, or larger
multimodal models.

The latest NVIDIA GR00T N1.7 documentation requires at least 16 GB of GPU
memory for inference. The current Base Router has 8 GB. Robotics should
therefore begin with small arms and SmolVLA-class models, while GR00T is a
future hardware-tier opportunity.

#### Model and capability map

Fit labels:

- **Ship:** Strong current product fit
- **Pilot:** Technically credible, but requires product testing and packaging
- **Future:** Requires different hardware, restricted access, or more maturity

| Capability | Current candidates | Practical fit | Product | Status |
|---|---|---|---|---|
| Text and document redaction | Microsoft Presidio with GLiNER PII or another local NER model | Rules plus a small entity model can detect, mask, replace, or encrypt personal information locally | All products, strongest on Base Router | Ship |
| Image redaction | Presidio Image Redactor plus local OCR | Finds text in an image, detects personal information, and covers or replaces it | Base Drive, Base Router | Pilot |
| Prompt attack detection | Llama Prompt Guard 2, 22M or 86M | Very small classifier for prompt injection and jailbreak attempts | All products | Ship |
| Content safety | Qwen3Guard 0.6B, 4B, or 8B | The 0.6B version is the realistic default for low-memory products | All products | Ship |
| Policy and groundedness judging | Granite Guardian 4.1 8B | Useful for custom rules, hallucination checks, and RAG evaluation, but tight on an 8 GB device | Capable laptops, future higher-memory router | Pilot |
| Basic OCR | PP-OCRv6 tiny, small, and medium | 1.5M to 34.5M parameter tiers make OCR viable even on modest hardware | All products | Ship |
| Document parsing | PaddleOCR-VL-1.6 0.9B, Docling | Reads text, layout, tables, formulas, charts, and document structure | Base Drive, Base Router | Ship |
| Local file search | EmbeddingGemma 308M | Multilingual embeddings for semantic search, classification, clustering, and RAG | All products | Ship |
| Text translation | NLLB-200 600M or 1.3B; smaller language-pair models | Broad offline language coverage, with smaller models chosen for lower-memory devices | Base Drive, Base Router | Ship |
| Speech translation | SeamlessM4T Medium 1.2B or Large v2 2.3B | One model can support speech recognition and speech/text translation across many languages | Base Drive, Base Router | Pilot |
| Speech transcription | Moonshine 27M or 61M; Whisper tiny, base, or small | Moonshine is attractive for low-cost live voice; Whisper offers a mature multilingual ladder | All products | Ship |
| Voice interface plumbing | Silero VAD, openWakeWord, Piper | Lightweight local speech detection, opt-in wake phrase, and text-to-speech | Router products | Pilot |
| Object detection | YOLO26 nano/small or YOLOE variants | Real-time detection and tracking on edge GPUs; commercial licensing must be reviewed | Base Router | Pilot |
| Image segmentation | MobileSAM, YOLO26-seg, SAM 2.1 | MobileSAM is about 39 MB; SAM 2.1 is better suited to capable laptops or GPU devices | Base Drive, Base Router | Pilot |
| Image and video understanding | Small LFM2-VL, Gemma, or Qwen multimodal variants | Useful for describing scenes, inspecting images, and combining vision with text | Capable laptops, Base Router | Pilot |
| General local assistant | LFM2.5 1.2B; Gemma 4 E2B/E4B; Qwen3.5 small models; Granite 4.0 Micro | Small general models coordinate specialist results and handle writing, extraction, and tool use | All products, size matched to hardware | Ship |
| Lightweight robotics | SmolVLA 450M | Can run on consumer hardware and is designed for affordable robot arms after task-specific adaptation | Base Drive Pro, Base Router | Pilot |
| General-purpose robotics | GR00T N1.7; Gemini Robotics On-Device | GR00T N1.7 needs 16 GB or more GPU memory; Gemini access is restricted | Future Base Edge or Robotics product | Future |

#### Recommended use-case packs

##### 1. Private Document Desk

**Models and tools:** PP-OCRv6, PaddleOCR-VL or Docling, EmbeddingGemma,
Presidio, GLiNER PII, and a small general model.

**What the buyer does:**

- Drops in PDFs, scans, forms, spreadsheets, or photos
- Searches across them in ordinary language
- Extracts tables, names, dates, and action items
- Creates summaries and drafts
- Produces a redacted copy before sharing

**Best products:** Base Drive, Base Router Lite, Base Router

**Positioning language:**

> Search, summarize, and redact your documents without sending the originals
> away.

This is the strongest near-term professional use case because it combines the
existing privacy story with specialist models that fit the current hardware.

##### 2. Local Privacy Gateway

**Models and tools:** Presidio, GLiNER PII, Llama Prompt Guard 2, Qwen3Guard,
and the existing consent log.

**What the buyer does:**

- Uses local AI normally
- Lets the Base identify sensitive fields automatically
- Reviews what would be removed before outside escalation
- Sends a sanitized request only after approval
- Keeps a local record of what left and why

**Best products:** Base Router Lite and Base Router

**Positioning language:**

> When outside help is needed, your Base can remove names, account numbers,
> and other private details before anything leaves.

This makes the current consent feature materially more valuable. Consent alone
answers "may this leave?" Redaction also answers "what actually needs to
leave?"

##### 3. Offline Language Desk

**Models and tools:** NLLB-200, SeamlessM4T, Whisper or Moonshine, and Piper.

**What the buyer does:**

- Translates documents without uploading them
- Transcribes and translates interviews or meetings
- Creates bilingual drafts
- Uses translation in locations with weak or no connectivity

**Best products:** Base Drive and Base Router

**Positioning language:**

> Translate documents and conversations where the cloud cannot reach, or
> where it should not be invited.

This is especially strong for travel, field work, schools, community
organizations, legal intake, and multilingual small businesses.

##### 4. Private Voice Notes

**Models and tools:** Moonshine or Whisper, Silero VAD, optional openWakeWord,
Piper, and a small general model.

**What the buyer does:**

- Dictates notes
- Transcribes meetings and interviews
- Creates summaries and follow-up lists
- Uses read-aloud and accessibility features
- Keeps raw recordings and transcripts local

**Best products:** Base Drive and Base Router

**Positioning language:**

> Turn speech into notes without sending the recording to a transcription
> service.

The box should not include an always-listening microphone by default. Voice
input should come from an explicitly approved phone, computer, or USB
microphone, with a visible recording indicator.

##### 5. Local Vision Workshop

**Models and tools:** MobileSAM, YOLO26 or YOLOE, SAM 2.1, and a small
vision-language model.

**What the buyer does:**

- Counts objects or inventory
- Marks the exact area of damage or a defect
- Removes image backgrounds
- Sorts field images
- Tracks selected objects in recorded video
- Inspects a workspace without uploading camera footage

**Best products:** Base Drive Pro and Base Router

**Positioning language:**

> Inspect, count, and mark what a camera sees while the footage stays on your
> equipment.

Offline Base should not imply that the router contains a hidden camera. The
camera is a separately approved input, and local processing is the feature.

##### 6. Local Robotics Lab

**Models and tools:** SmolVLA, LeRobot, small object detection and
segmentation models.

**What the buyer does:**

- Connects a low-cost robot arm
- Records demonstrations
- Adapts a small action model to a narrow task
- Runs perception and control without an internet dependency

**Best products:** Base Drive Pro and Base Router

**Positioning language:**

> A private, local starting point for teaching a small robot a specific job.

Do not claim general-purpose robot autonomy. SmolVLA still requires compatible
hardware, demonstrations, testing, and physical safety controls. Latest
GR00T-class support belongs to a future higher-memory product.

#### Audience personas

These personas are intentionally narrower than "students," "white-collar
workers," and "old people." Broad demographic labels do not identify a product
to build. Each primary persona below has a repeated job, a likely buyer, a
clear product fit, and a small first release.

##### 1. Maya, the commuter college student

**Profile**

- 20 years old
- Attends a public college and commutes by train
- Uses a midrange laptop with 8 to 16 GB of memory
- Works part time and studies in places with inconsistent internet
- Comfortable with chat apps but does not want to configure models
- Price sensitive and likely to buy with help from a parent

**What she is trying to accomplish**

- Understand a difficult reading without replacing the reading
- Search lecture slides, class notes, and assigned PDFs
- Turn recorded lectures or personal voice notes into study material
- Build practice questions, flashcards, and outlines
- Improve a draft while keeping her own voice
- Translate a passage or explain unfamiliar academic language
- Continue working during a commute or campus network outage

**Current frustration**

Her files, notes, recordings, and conversations are split across several apps.
Useful AI tools often require another subscription, upload class material, or
stop working when connectivity is poor. She also cannot easily tell whether a
chatbot's answer came from her assigned material or from a guess.

**Best product**

**Base Drive** with a **Study Pack**. Base Drive Lite can be the budget option,
but the standard Base Drive is the better default because document search and
transcription benefit from more capable laptops.

**Study Pack MVP**

1. Add a course folder containing PDFs, slides, notes, and images.
2. Ask questions with page-level citations back to the source.
3. Create a study guide, practice quiz, or flashcards from selected files.
4. Transcribe a recording and link the transcript to the course.
5. Explain or translate selected text.
6. Review a draft with modes such as "clarity," "structure," and "quiz me."
7. Work with Wi-Fi turned off.

**Day-one workflow**

> Maya plugs in Base Drive, selects "Set up for school," adds her biology
> folder, and asks for ten practice questions based only on this week's lecture
> and assigned chapter. On the train, she answers them offline and asks for an
> explanation with the exact page that supports it.

**UX requirements**

- Course-based folders instead of model selection
- Visible citations and a "use only my files" switch
- Clear separation between tutoring and writing the assignment for her
- Fast import from ordinary folders
- Storage controls for deleting a class after the term
- A simple laptop compatibility check before purchase

**Trust and safety requirements**

- Do not market it as a way to evade academic integrity rules.
- Label unsupported answers instead of inventing a citation.
- Make recording consent clear before lecture transcription.
- Keep student files and transcripts local by default.

**Positioning line**

> Your classes, notes, and study help in one private place, even when the
> internet is not.

##### 2. Daniel, the document-heavy knowledge worker

**Profile**

- 38 years old
- Works in operations, consulting, finance, insurance, HR, or account
  management
- Spends much of the day reading documents, preparing meetings, and writing
  follow-up material
- Handles internal, client, employee, or financial information
- Understands the benefit of AI but hesitates to upload work files to a public
  chatbot
- May buy personally, expense it, or recommend it to a small team

**What he is trying to accomplish**

- Search across project files, policies, reports, and meeting notes
- Summarize long documents and compare versions
- Turn meetings into decisions, owners, and next steps
- Draft emails, briefs, reports, and client updates from approved material
- Extract names, dates, totals, obligations, and risks
- Redact private details before sharing a document or asking an outside model
- Reuse repeatable workflows without copying the same instructions every time

**Current frustration**

The general chatbot is helpful, but company policy or client expectations may
prevent him from using it with the work that matters most. Enterprise AI can
require procurement, per-seat subscriptions, and another system to administer.
He needs a tool that works with files he already has and makes its privacy
boundary obvious.

**Best product**

**Base Router** with a **Private Work Desk** for an office or household with
multiple devices. **Base Drive Pro** is the best individual or traveling
option.

**Private Work Desk MVP**

1. Create a private workspace from local folders.
2. Search and answer with source citations.
3. Compare two documents and list meaningful changes.
4. Transcribe a meeting and produce decisions, owners, and follow-ups.
5. Extract structured information into a table.
6. Draft from selected sources and saved company templates.
7. Detect and review sensitive details before export or outside escalation.
8. Keep a local activity record of exports and approved outside requests.

**Day-one workflow**

> Daniel adds a project folder before a client meeting. Base Router produces a
> briefing from the contract, recent notes, and status report. After the
> meeting it turns the recording into decisions and follow-ups. When Daniel
> asks for outside help with a difficult rewrite, the Base shows which client
> names and account details it will remove before asking for approval.

**UX requirements**

- Workspaces named after clients, projects, or departments
- Source citations and an obvious local-only status
- Reusable actions such as "meeting brief," "compare," and "redact"
- Review screens before files are exported or information leaves
- Role-based access for shared routers
- Plain-language records that a manager can inspect

**Trust and safety requirements**

- Do not claim that local processing automatically satisfies a profession's
  legal or regulatory obligations.
- Require human review for redaction, financial extraction, and final advice.
- Keep source permissions aligned with user and team access.
- Clearly distinguish measured local capability from optional outside help.

**Positioning line**

> Use AI on the documents your work depends on without handing those documents
> to another company.

##### 3. Linda, the independent older adult

**Profile**

- 72 years old
- Lives independently and uses a phone and laptop for email, bills, photos,
  travel, and family communication
- Can use familiar apps but becomes frustrated by account creation, changing
  interfaces, small text, and technical setup
- Values privacy and does not want every question or family document stored by
  an online service
- May buy the product herself, but an adult child may purchase or set it up

**What she is trying to accomplish**

- Understand a confusing letter, bill, insurance notice, or form
- Have a document or message read aloud
- Draft a clear reply without navigating a complex editor
- Translate a message or travel document
- Search family records, recipes, notes, and scanned photographs
- Turn spoken thoughts into a note or letter
- Check a suspicious message for warning signs before responding
- Get help without remembering another password or paying a monthly fee

**Current frustration**

Many digital assistants are designed around small controls, hidden settings,
frequent changes, and accounts she does not understand. Sending financial,
health, or family documents to an unknown cloud service feels unsafe. When
something goes wrong, recovery instructions are often harder than the original
task.

**Best product**

**Base Router Lite** with an **Everyday Help Pack**. It is always available at
home, works from devices she already knows, and can be set up once by the buyer
or a trusted family helper. Base Drive is less suitable because it introduces
an installation and laptop compatibility decision.

**Everyday Help Pack MVP**

1. Four large starting actions: "Explain a document," "Write a reply," "Read
   this aloud," and "Ask about my files."
2. Take a photo of a letter or open an existing file.
3. Explain it in plain language while showing the original text.
4. Read selected text aloud at an adjustable speed.
5. Accept dictation for notes and replies.
6. Flag common scam signals without declaring a message safe.
7. Provide a trusted-helper setup and recovery mode.
8. Work without an always-listening microphone or camera.

**Day-one workflow**

> Linda opens Offline Base from a bookmark on her tablet and photographs an
> unfamiliar insurance letter. It reads the letter aloud, explains each
> section in plain language, and highlights the phone number and deadline. It
> does not tell her what medical or financial decision to make. She can send
> the original and explanation to her daughter if she chooses.

**UX requirements**

- Large text, strong contrast, generous spacing, and large touch targets
- One question or decision per screen
- Familiar verbs instead of AI terms
- Optional voice input with an unmistakable recording indicator
- A persistent "show me the original" action
- Undo, confirmation, and a simple way back home
- Printed setup and recovery instructions
- A helper role that cannot silently read private conversations

**Trust and safety requirements**

- Never market the product as a replacement for a doctor, lawyer, financial
  adviser, emergency service, or human caregiver.
- Treat scam detection as a warning aid, not a guarantee.
- Avoid implying that older adults are incapable or share the same needs.
- Require explicit approval before a helper can access files or settings.
- Make deletion, microphone status, and outside escalation easy to understand.

**Positioning line**

> Everyday help with letters, forms, and family files, kept in your home and
> made simple to use.

#### Persona build order

| Order | Persona | Why build now | Initial product |
|---|---|---|---|
| 1 | Daniel, knowledge worker | Highest willingness to pay, strongest privacy pain, and the document stack overlaps with the professional market already present in the site | Base Router plus Private Work Desk |
| 2 | Maya, student | Large audience and a natural Base Drive fit; reuses document search, transcription, translation, and citation work from the professional product | Base Drive plus Study Pack |
| 3 | Linda, older adult | Strong differentiation and social value, but requires the most accessibility testing, support design, and careful safety language | Base Router Lite plus Everyday Help Pack |

This order is for product development, not market importance. The first two
personas build the document, search, transcription, and redaction foundation.
The older-adult experience can then simplify those proven capabilities instead
of exposing an unfinished general chat product behind larger buttons.

#### Shared product foundation

The three personas should not become three unrelated applications. They share
the same underlying capabilities:

- Local document import, OCR, and search
- Answers tied to source material
- Speech transcription and read-aloud
- Translation
- Redaction and consent before outside help
- Simple task templates
- Clear storage and deletion controls

The difference is the starting experience:

- **Student:** course folders, study actions, and citations
- **Knowledge worker:** project workspaces, extraction, redaction, and records
- **Older adult:** a few large everyday actions, read-aloud, and trusted setup

#### Vertical positioning

##### Law firms

Lead with:

- Local OCR and search over matter files
- Privilege-conscious redaction before sharing
- Deposition and interview transcription
- Translation for multilingual intake
- A local record of every outside escalation

Avoid claiming that model output is legal advice or that automated redaction is
infallible.

##### Clinics

Lead with:

- Local intake transcription
- Search and summarization over approved records
- Removal of identifiers before an outside request
- Translation support for patient communication

Avoid claiming that local processing alone creates HIPAA compliance. Human
review is required for clinical decisions and redaction quality.

##### Schools and districts

Lead with:

- OCR and search across curriculum and district documents
- Local translation for family communication
- Student-information redaction
- Private speech-to-text accessibility tools
- Prompt and content safeguards

Avoid claiming that a safety classifier replaces district policy or staff
review.

##### Field teams

Lead with:

- Offline translation and transcription
- OCR from signs, labels, and forms
- Image segmentation and object counting
- Search over local manuals
- Operation without a reliable connection

##### Workshops and small manufacturers

Lead with:

- Visual inspection and object counting
- Local work-instruction search
- Voice notes with hands occupied
- Narrow robotics experiments

Avoid promising production quality control until the model has been validated
on the customer's own camera, lighting, objects, and failure conditions.

#### Recommended model-routing architecture

```text
User or approved device
        |
        v
Local input checks
  - prompt attack detection
  - personal information detection
  - file and media type detection
        |
        v
Task router
  - documents -> OCR / document parser
  - search -> embedding model
  - speech -> VAD / transcription / translation
  - images -> detection / segmentation / vision model
  - robot -> perception / task-specific action model
        |
        v
Small general model
  - combines results
  - explains findings
  - drafts the final response
        |
        v
Can the task be completed locally?
  - yes -> return result
  - no  -> redact locally -> show consent prompt -> optional outside model
        |
        v
Local activity and approval record
```

This architecture gives the name **Base Router** a product meaning beyond
network hardware.

The detailed memory, retrieval, redaction, and portability design is in
[Private Memory Router Architecture](./PRIVATE_MEMORY_ROUTER_ARCHITECTURE.md).

#### Priority order

##### Priority 0: Redaction, OCR, and local search

Build first:

- Presidio plus a selected PII model
- PP-OCRv6
- Docling or PaddleOCR-VL
- EmbeddingGemma
- One small general model

Why:

- Fits every current product
- Directly supports law, clinic, school, home, and field positioning
- Strengthens the existing consent story
- Can be tested with ordinary files rather than new hardware

##### Priority 1: Speech and translation

Build next:

- Moonshine and Whisper hardware comparison
- NLLB-200 text translation
- SeamlessM4T on capable hardware
- Piper for local read-aloud

Why:

- Clear offline value
- Strong accessibility and field use cases
- Works through existing phones and computers without adding a hidden
  microphone to the box

##### Priority 2: Vision

Pilot:

- MobileSAM
- YOLO26 nano or a commercially suitable alternative
- SAM 2.1 on capable laptops and the Jetson

Why:

- Strong visual demo
- Good fit for the Jetson
- Requires camera consent, licensing review, and task-specific evaluation

##### Priority 3: Robotics

Pilot only:

- SmolVLA with a supported low-cost arm
- A bounded pick, place, sort, or inspection task

Defer:

- GR00T N1.7
- Humanoid and general-purpose robot claims
- Safety-critical autonomous control

Why:

- The current 8 GB Jetson is below GR00T N1.7's documented 16 GB minimum
- Robotics requires a complete hardware, data, safety, and support product,
  not only a model download

#### Model catalog guidance

The repo currently includes fast-changing model names in
`app/lib/drive-config.ts` and `app/lib/throughput.ts`. Several names are now
real current releases, including Gemma 4, Qwen3.5, Qwen3.6, and LFM2.5.
However:

- Exact checkpoints, licenses, quantizations, and device performance still
  need verification.
- IBM's currently documented general family is Granite 4.0, while the repo
  says Granite 4.1. Granite Guardian 4.1 is documented, but the exact general
  Granite 4.1 checkpoint should be verified before a public claim.
- A model existing is not proof that it runs well on the advertised product.
- Parameter count does not equal memory use, speed, or output quality.

Public pages should therefore say what the product can do and reserve exact
model names for a dated compatibility page or technical appendix.

#### Evaluation gate before any public claim

Every model pack should pass the same gate:

1. Confirm model license and commercial distribution rights.
2. Record exact checkpoint, revision, quantization, and runtime.
3. Measure download size, peak memory, cold start, latency, and sustained
   throughput on each product.
4. Test accuracy on representative customer files, languages, images, or
   recordings.
5. Document known failure modes and when human review is required.
6. Verify that the feature works with networking disabled.
7. Confirm that logs, temporary files, embeddings, and derived media remain
   local.
8. Test the redaction and consent path before any outside escalation.
9. Write buyer-facing language from the measured result, not the model card.

#### Claims to avoid

Do not claim:

- Perfect redaction or guaranteed removal of all personal information
- Certified legal, medical, educational, or translation accuracy
- HIPAA or FERPA compliance solely because processing is local
- Real-time vision on every Base product
- That an 8 GB Jetson can run the latest GR00T model
- General-purpose robot autonomy
- That every listed model is loaded or running at the same time
- That open weights automatically permit every commercial use

Prefer:

- "Designed to keep this work local"
- "Helps identify and remove sensitive details"
- "Tested on this Base product with this model pack"
- "Human review is recommended before sharing or acting"
- "Outside help is optional and requires approval"

#### Source notes

Primary and project sources used for this research:

- [Microsoft Presidio](https://microsoft.github.io/presidio/) supports local
  detection and anonymization of personal information in text and images.
- [GLiNER Multi PII](https://huggingface.co/urchade/gliner_multi_pii-v1)
  provides a small multilingual entity model intended for PII detection.
- [Meta NLLB-200](https://ai.meta.com/research/publications/no-language-left-behind-scaling-human-centered-machine-translation/)
  covers 200 languages and more than 40,000 translation directions.
- [Meta SeamlessM4T](https://github.com/facebookresearch/seamless_communication)
  supports text and speech translation, transcription, and speech output.
- [Meta Omnilingual MT](https://ai.meta.com/research/publications/omnilingual-mt-machine-translation-for-1600-languages/)
  is a March 17, 2026 research release covering more than 1,600 languages. It
  belongs on the watchlist until a supported local deployment path is clear.
- [PaddleOCR](https://paddlepaddle.github.io/PaddleOCR/main/en/index.html)
  released PaddleOCR-VL-1.6 on May 28, 2026 and PP-OCRv6 on June 11, 2026.
- [Docling](https://github.com/docling-project/docling) provides local parsing
  for PDFs, office files, images, audio, email, tables, formulas, and layout.
- [EmbeddingGemma](https://ai.google.dev/gemma/docs/embeddinggemma) is a 308M
  multilingual embedding model optimized for everyday devices.
- [OpenAI Whisper](https://github.com/openai/whisper) provides multilingual
  transcription and translation models from 39M to 1.55B parameters.
- [Moonshine](https://github.com/moonshine-ai/moonshine) targets low-latency
  local speech applications, including transcription and command recognition.
- [Silero VAD](https://github.com/snakers4/silero-vad) is a lightweight local
  voice activity detector.
- [Piper](https://github.com/OHF-Voice/piper1-gpl) is a local text-to-speech
  engine. Its GPL license requires product-level review.
- [SAM 2.1](https://github.com/facebookresearch/sam2) supports image and video
  segmentation under Apache 2.0 model and code licensing.
- [MobileSAM](https://github.com/ChaoningZhang/MobileSAM) provides a lightweight
  Segment Anything variant with an approximately 39 MB checkpoint.
- [Ultralytics YOLO26](https://docs.ultralytics.com/models/yolo26) supports
  detection, segmentation, pose, classification, and oriented boxes. Review
  [Ultralytics licensing](https://www.ultralytics.com/license) before
  distribution.
- [Qwen3Guard](https://qwenlm.github.io/blog/qwen3guard/) provides 0.6B, 4B,
  and 8B safety models.
- [Llama Prompt Guard 2](https://huggingface.co/meta-llama/Llama-Prompt-Guard-2-86M)
  provides 22M and 86M prompt attack classifiers.
- [Granite Guardian 4.1](https://www.ibm.com/granite/docs/models/guardian)
  supports safety, groundedness, tool-call hallucination, and custom criteria.
- [Gemma 4](https://ai.google.dev/gemma/docs/core/model_card_4) provides E2B,
  E4B, 12B, 26B A4B, and 31B general model tiers.
- [Liquid LFM2.5](https://www.liquid.ai/blog/introducing-lfm2-5-the-next-generation-of-on-device-ai)
  includes a 1.2B model designed for on-device assistants and tool use.
- [SmolVLA](https://huggingface.co/blog/smolvla) is a 450M open robotics model
  designed to run on consumer hardware.
- [NVIDIA GR00T N1.7](https://github.com/NVIDIA/Isaac-GR00T) is the latest
  early-access GR00T release and documents a 16 GB or greater GPU memory
  requirement for inference.
- [Gemini Robotics On-Device](https://deepmind.google/blog/gemini-robotics-on-device-brings-ai-to-local-robotic-devices/)
  is optimized for local robot control but remains a restricted trusted-tester
  offering.
- [NVIDIA Jetson Orin Nano Super](https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/nano-super-developer-kit/)
  provides 8 GB memory, up to 67 INT8 TOPS, and a 7 to 25 watt power range.
- [Raspberry Pi 5](https://www.raspberrypi.com/products/raspberry-pi-5/) uses a
  2.4 GHz quad-core Arm Cortex-A76 CPU and is available in several memory tiers.

<!-- source-document-end -->

## Source Document: Local AI Agents Edge Deployment Research

Source file: `docs/local-ai-agents-edge-deployment-research.md`

<!-- source-document-start -->

### Local AI Agents on Raspberry Pi, NVIDIA Jetson, and Portable Drives

Research snapshot: June 13, 2026

This document evaluates local AI agent frameworks, inference runtimes, small
language models, and deployment patterns for:

1. A standalone Raspberry Pi.
2. A standalone NVIDIA Jetson.
3. A portable flash drive or external SSD used with another computer.
4. A combined system where the Pi, Jetson, drive, and optional laptop each do
   the work they are best suited to do.

The recommendations are based primarily on official project repositories,
documentation, model cards, and hardware documentation. Fast-moving projects
must still be pinned and tested on the exact target image before shipping.

#### Executive Recommendation

There is no single agent that is best on every device.

The strongest general architecture is:

- **Raspberry Pi:** always-on gateway, scheduler, messaging bridge, policy
  enforcement, and lightweight agent process.
- **Jetson:** model inference server using `llama.cpp`, Ollama, or vLLM.
- **External USB SSD or NVMe drive:** models, retrieval documents, checksums,
  offline installers, configuration templates, licenses, and backups.
- **Laptop or workstation:** optional Cline, Pi, Aider, OpenCode, or another
  interactive coding client connected to the Jetson model endpoint.

For a single-device product:

| Device | Recommended baseline | Realistic local model class |
| --- | --- | --- |
| Raspberry Pi 5, 8 GB | `llama.cpp` plus Pi, ZeroClaw, PicoClaw, or nanobot | 1B to 3B, Q4, 8K context |
| Raspberry Pi 5, 16 GB | `llama.cpp` plus Pi, Aider, ZeroClaw, or nanobot | 3B to 8B, Q4, usually 8K to 16K context |
| Jetson Orin Nano, 8 GB | Ollama or CUDA `llama.cpp` plus Pi, Aider, ZeroClaw, or nanobot | 2B to 4B comfortably; 7B is tight |
| Jetson Orin NX, 16 GB | Ollama, CUDA `llama.cpp`, or Jetson AI Lab container | 7B to 14B, Q4 |
| Jetson AGX Orin, 32 GB | CUDA `llama.cpp`, Ollama, or vLLM | 14B to 27B or 32B, depending on context |
| Jetson AGX Orin, 64 GB | CUDA `llama.cpp`, Ollama, or vLLM | 32B comfortably; some 70B Q4 deployments |
| Jetson Thor, 128 GB | vLLM or optimized NVIDIA stack | 70B class and multi-model workloads |

These are conservative planning ranges, not guarantees. Model architecture,
quantization, context length, KV cache type, runtime, GPU offload, and agent
overhead all affect fit. A model fitting in memory does not mean a long-context
agent will run reliably.

#### Terms That Must Stay Separate

An edge AI system has four independent layers:

1. **Agent:** decides what to do, manages messages and memory, and calls tools.
   Examples: OpenClaw, Hermes Agent, Pi, Cline, Aider, and ZeroClaw.
2. **Inference runtime:** loads the model and exposes a local API. Examples:
   `llama.cpp`, Ollama, vLLM, and LocalAI.
3. **Model:** generates text and tool calls. Examples: Qwen3.5 4B,
   Ministral 3 3B, and Phi-4 Mini.
4. **Storage or boot media:** contains models and software. A flash drive is
   not a compute device and cannot run an agent without a compatible host.

Most agent projects do not run a model themselves. They call an inference
runtime through an OpenAI-compatible or provider-specific API.

#### Agent Comparison

##### Summary Matrix

| Agent | Primary use | Local model support | Edge suitability | Main concern |
| --- | --- | --- | --- | --- |
| LocalClaw | Local-first personal assistant | Ollama, LM Studio, vLLM, OpenAI-compatible | Experimental Pi/Jetson option | Small project and slower maintenance cadence |
| OpenClaw | Always-on personal assistant and channel gateway | Ollama and custom providers | Pi gateway, larger Jetson inference | Official Pi guidance assumes remote inference |
| Hermes Agent | Persistent general assistant with skills and tools | Ollama, vLLM, `llama.cpp`, custom endpoint | Larger Jetson or split deployment | Large context expectation and heavier dependencies |
| Pi | Minimal terminal coding-agent harness | Ollama, vLLM, LM Studio, OpenAI-compatible | Strong edge coding choice | No built-in sandbox |
| Cline | IDE and CLI coding agent | Ollama and LM Studio | Laptop client to Jetson | High context and memory demand |
| Aider | Terminal pair programmer | Ollama and OpenAI-compatible routes | Strong edge coding choice | Coding focused, not a personal gateway |
| OpenCode | Terminal coding agent | Ollama, LM Studio, `llama.cpp` | Jetson or laptop client | ARM64 compatibility regressions require testing |
| Goose | General and coding agent with extensions | Ollama and other local providers | Jetson; Pi client is plausible | More runtime overhead than minimal agents |
| Qwen Code | Terminal coding agent | Ollama, vLLM, OpenAI-compatible | Jetson or laptop client | Node process and context overhead |
| nanobot | Lightweight personal agent | Ollama, vLLM, OpenAI-compatible | Good Pi/Jetson candidate | Younger ecosystem than OpenClaw |
| ZeroClaw | Lightweight Rust personal agent | Ollama and other providers | Strong Pi gateway candidate | Newer project, smaller operational history |
| PicoClaw | Very small Go personal agent | Ollama and compatible endpoints | Strong low-resource experiment | Project performance claims need independent tests |

##### LocalClaw

[LocalClaw](https://github.com/sunkencity999/localclaw) is a local-first fork
and adaptation of OpenClaw. It is designed around small local models and
constrained context windows.

Useful properties:

- Supports Ollama, LM Studio, vLLM, and OpenAI-compatible endpoints.
- Uses aggressive context pruning and shorter tool output.
- Includes repair behavior for models that emit tool calls as text instead of
  valid structured output.
- Keeps its state under `~/.localclaw`, allowing it to coexist with OpenClaw.
- Requires Node.js 22 or newer.

Fit:

- Promising for a Jetson with a 3B to 9B model.
- Potentially usable on a Pi 5 with a 1B to 3B model.
- Better aligned with short-context edge models than full OpenClaw or Hermes.

Risks:

- It had a much smaller maintenance footprint than the major frameworks at the
  research snapshot.
- Raspberry Pi and Jetson support are not documented as first-class deployment
  targets.
- Text repair makes weaker models usable, but it does not make their actions
  reliable or secure.

Recommendation: retain as an experimental local-first option, but do not make it
the only production agent until the selected model passes the tool-call and
recovery tests in this document.

##### OpenClaw

[OpenClaw](https://github.com/openclaw/openclaw) is an always-on personal
assistant gateway with channels, persistent sessions, skills, tools, and
multiple agents. It supports [Ollama](https://docs.openclaw.ai/providers/ollama)
and other local or remote providers.

Useful properties:

- Mature personal-assistant and messaging architecture.
- Persistent file-based workspace and memory.
- Broad channel and skill ecosystem.
- One gateway can manage multiple agent profiles.

Hardware reality:

- The official
  [Raspberry Pi guide](https://docs.openclaw.ai/install/raspberry-pi) treats the
  Pi primarily as the gateway and recommends remote model inference.
- Its
  [local-model guidance](https://docs.openclaw.ai/gateway/local-models) warns
  that local models substantially increase hardware, context, and security
  requirements.
- Its Ollama guidance recommends large context windows, which are expensive on
  8 GB and 16 GB devices.

Fit:

- Good on a Pi when the model is hosted by a Jetson.
- Plausible as an all-in-one system on AGX Orin 32 GB or larger.
- Not a sensible all-in-one target for a Pi 5 with a small model unless the
  tools and prompt are heavily reduced.

Recommendation: use OpenClaw for the gateway and channel layer, not as the
default same-device Pi inference stack.

##### Hermes Agent

[Hermes Agent](https://github.com/NousResearch/hermes-agent) is a persistent
general assistant with memory, skills, tool use, a learning loop, messaging
channels, and several execution backends. It supports
[Ollama and self-hosted endpoints](https://hermes-agent.nousresearch.com/docs/integrations/providers).

Useful properties:

- Persistent memories and learned skills.
- Rich built-in tool set.
- Local, Docker, SSH, and other execution backends.
- Custom OpenAI-compatible endpoint support.
- Official [Ollama integration](https://docs.ollama.com/integrations/hermes).

Constraints:

- The official
  [Hermes FAQ](https://hermes-agent.nousresearch.com/docs/reference/faq)
  specifies a 64,000-token minimum for a local model. That makes a small
  same-device deployment expensive.
- Installation requires a larger Python, Node.js, media, and utility dependency
  set than a minimal edge agent.
- Its public issue history contains
  [Jetson ARM64 installation problems](https://github.com/NousResearch/hermes-agent/issues/38580)
  and an explicit request for a
  [supported air-gapped installation process](https://github.com/NousResearch/hermes-agent/issues/17696).
  This means offline use after installation is different from offline
  installation.

Fit:

- Good on AGX Orin 32 GB, 64 GB, or Jetson Thor.
- Good when the agent runs on a Pi or another host and calls a large Jetson
  model.
- Poor fit for a Pi 5 running both Hermes and a small local model.

Recommendation: offer Hermes as a high-capability profile for larger Jetsons,
not as the minimum hardware baseline.

##### Pi Coding Agent

[Pi](https://pi.dev/) and its
[current repository](https://github.com/earendil-works/pi) provide a minimal
terminal agent harness. This project name is unrelated to Raspberry Pi
hardware.

Useful properties:

- Small, inspectable agent loop.
- Built-in `read`, `write`, `edit`, and `bash` tools.
- Extensions, skills, prompt templates, themes, and packages.
- Supports Ollama, vLLM, LM Studio, and OpenAI-compatible endpoints.
- Avoids many high-overhead orchestration features by default.

Local provider example:

```json
{
  "providers": {
    "local": {
      "baseUrl": "http://127.0.0.1:8080/v1",
      "api": "openai-completions",
      "apiKey": "local",
      "compat": {
        "supportsDeveloperRole": false,
        "supportsReasoningEffort": false
      },
      "models": [
        {
          "id": "edge-model"
        }
      ]
    }
  }
}
```

Security constraint:

Pi does not provide a built-in sandbox. Its shell and file tools run with the
permissions of the current user. On an appliance, run it under an unprivileged
service account, restrict its workspace, and place it in a container or stronger
sandbox when it can act without direct supervision.

Fit:

- Best lightweight coding-agent candidate for Pi and Jetson.
- Particularly useful as a laptop client to a model running on Jetson.
- Not a complete messaging or personal-assistant gateway on its own.

##### Cline

[Cline](https://github.com/cline/cline) is a full IDE and CLI coding agent with
checkpoints, file editing, browser use, terminal tools, MCP integrations, and
review workflows.

Official
[local-model guidance](https://docs.cline.bot/running-models-locally/overview)
supports Ollama and LM Studio and recommends compact prompts. Current guidance
places typical small quantized setups in the 16 GB to 32 GB RAM range and
recommends a large context window for coding.

Fit:

- Good on a developer laptop connected to a Jetson endpoint.
- Plausible on a larger Jetson when a desktop environment and IDE are desired.
- Poor fit as the primary same-device agent on an 8 GB Pi or Jetson.

Recommendation: support Cline as a client integration, not as the core edge
appliance process.

##### Aider

[Aider](https://github.com/Aider-AI/aider) is a mature terminal pair programmer.
Its repository map helps limit context use, and its
[Ollama documentation](https://aider.chat/docs/llms/ollama.html) supports local
models through the `ollama_chat` route.

Fit:

- Strong coding option on Pi 5 16 GB or Jetson.
- Strong laptop client to a Jetson model.
- Lower operational complexity than a general personal assistant.

Constraint: Aider is for editing code repositories. It does not replace an
always-on assistant gateway, messaging hub, or general automation daemon.

##### OpenCode

[OpenCode](https://github.com/anomalyco/opencode) is a terminal coding agent
with official provider configurations for Ollama, LM Studio, and `llama.cpp`.

Fit:

- Attractive on x86 laptops and tested Jetson builds.
- Can use a remote Jetson endpoint from another device.

Risk: the project has had ARM64 issues including illegal-instruction failures on
[Raspberry Pi](https://github.com/anomalyco/opencode/issues/13268) and
[crashes related to ARM page-size environments](https://github.com/anomalyco/opencode/issues/12474).
Pin an exact release and execute it on every supported OS image before
inclusion.

Recommendation: secondary option until ARM64 qualification is automated.

##### Goose

[Goose](https://github.com/aaif-goose/goose) is a Rust-based general and coding
agent with CLI, desktop, API, extensions, MCP support, and many providers. Its
[provider documentation](https://goose-docs.ai/docs/getting-started/providers)
includes Ollama.

Fit:

- Good on Jetson.
- Potential Pi client for a remote Jetson model.
- ARM64 release artifacts make appliance packaging plausible.

Risk: it has more components than Pi or Aider, and ARM-specific extension
behavior still needs qualification.

##### Qwen Code

[Qwen Code](https://github.com/QwenLM/qwen-code) is a terminal coding agent
supporting OpenAI-compatible, Anthropic-compatible, and Gemini-compatible APIs.
Its repository documents local Ollama and vLLM configurations.

Fit:

- Good with a Qwen model on a Jetson or remote laptop.
- Less suitable for an 8 GB same-device deployment because the Node agent
  process and coding context consume material memory before model inference.

##### Lightweight Personal-Agent Candidates

###### nanobot

[nanobot](https://github.com/HKUDS/nanobot) is a compact Python personal agent
with channels, tools, memory, MCP, a web interface, model routing, and local
provider support.

Use it when OpenClaw or Hermes is too heavy but a conversational personal agent
is still needed.

###### ZeroClaw

[ZeroClaw](https://github.com/zeroclaw-labs/zeroclaw) is a lightweight Rust
agent runtime with providers, channels, tools, memory, and MCP support. Official
container artifacts include Linux ARM64.

It is one of the strongest Pi gateway candidates because it combines a small
runtime with more personal-agent features than Pi or Aider.

###### PicoClaw

[PicoClaw](https://github.com/sipeed/picoclaw) is a small Go agent targeting
ARM, RISC-V, MIPS, and x86 systems. It is a useful candidate for the smallest
gateway profile.

Its published memory and startup numbers are project claims. Benchmark the
specific build, provider, channels, and tools before using those numbers in
product requirements.

##### Frameworks Not Recommended as the First Edge Baseline

- **OpenHands:** designed for a much heavier coding sandbox and service stack.
- **Full multi-agent platforms:** orchestration overhead and context
  multiplication work against small-model reliability.
- **IDE-only agents on the appliance:** a headless Pi or Jetson should expose an
  API and let the user's laptop provide the IDE.
- **Cloud-first agents that happen to accept a local endpoint:** local model
  support alone does not establish offline installation, offline operation, or
  ARM64 compatibility.

#### Inference Runtime Comparison

##### Recommended Runtime Matrix

| Runtime | Pi | Jetson | Portable drive host | Best use |
| --- | --- | --- | --- | --- |
| `llama.cpp` | Best baseline | Strong baseline | Best portable baseline | Minimal, direct GGUF serving |
| Ollama | Convenient if supported image is pinned | Best convenience option | Good after per-OS install | Model management and simple API |
| vLLM | Not recommended | AGX/Thor and throughput workloads | GPU workstation only | Concurrent serving and large models |
| LocalAI | Possible but heavier | Possible | Good compatibility layer | Multiple media/model backends |
| LM Studio | Not an appliance baseline | Not a headless baseline | Good desktop option | GUI-based local model use |

##### `llama.cpp`

[llama.cpp](https://github.com/ggml-org/llama.cpp) is the most portable baseline.
It uses GGUF model files, supports CPU inference, CUDA and other GPU backends,
and includes
[`llama-server`](https://github.com/ggml-org/llama.cpp/tree/master/tools/server),
which exposes OpenAI-compatible APIs.

Why it should be the compatibility baseline:

- Builds on ARM64 Linux.
- Does not require a separate model manager.
- The same GGUF file can be stored once and used by compatible x86 and ARM
  builds.
- Context length, GPU layers, parallel requests, and cache behavior are
  explicit.
- A portable drive can contain per-architecture binaries plus shared GGUF
  files.

Example Pi server:

```bash
./llama-server \
  -m /mnt/base-drive/models/model-q4_k_m.gguf \
  --host 127.0.0.1 \
  --port 8080 \
  -c 8192 \
  -np 1
```

Example Jetson server with CUDA offload:

```bash
./llama-server \
  -m /mnt/base-drive/models/model-q4_k_m.gguf \
  --host 192.168.50.10 \
  --port 8080 \
  -c 16384 \
  -np 1 \
  --n-gpu-layers 999
```

The actual binary path and GPU option support depend on how the selected
release was built.

##### Ollama

[Ollama](https://ollama.com/) is the easiest operational model manager and API
for many users. Its
[development documentation](https://docs.ollama.com/development) includes
JetPack 5 and JetPack 6 build targets, and NVIDIA's
[Jetson AI Lab guide](https://www.jetson-ai-lab.com/tutorials/ollama/) provides
Jetson-specific deployment guidance.

Recommended constrained-device settings:

```ini
OLLAMA_MODELS=/mnt/base-drive/ollama
OLLAMA_CONTEXT_LENGTH=8192
OLLAMA_MAX_LOADED_MODELS=1
OLLAMA_NUM_PARALLEL=1
```

The [Ollama FAQ](https://docs.ollama.com/faq) explains that parallel requests
increase context memory consumption. On an 8 GB or 16 GB appliance, use one
loaded model and one generation at a time unless a device benchmark proves more
capacity.

For LAN access, do not expose an unauthenticated endpoint to an untrusted
network. Bind to a private interface, filter by source subnet, or place an
authenticated reverse proxy in front of it.

##### vLLM

[vLLM](https://github.com/vllm-project/vllm) is appropriate when throughput,
continuous batching, and multiple clients matter more than minimal footprint.

Use it on:

- Jetson AGX Orin with sufficient memory.
- Jetson Thor.
- A larger NVIDIA workstation or server.

Do not make it the Raspberry Pi baseline. It adds packaging and memory
complexity without helping a single slow CPU-bound request.

##### LocalAI

[LocalAI](https://github.com/mudler/LocalAI) provides an OpenAI-compatible layer
over multiple model and media backends. It is useful when one appliance must
serve text, speech, images, and embeddings through a common API.

It is broader and heavier than a direct `llama.cpp` server. Use it only when the
additional backends are part of the product requirements.

#### Model Selection

##### What an Agent Model Must Do

Small-model selection cannot be based only on chat quality or benchmark score.
An agent model must:

1. Follow the exact chat template supported by the runtime.
2. Emit valid tool names and JSON arguments.
3. consume tool results across multiple turns.
4. Recover from tool errors without inventing success.
5. Resist instructions embedded in untrusted documents.
6. Stay coherent after context pruning.
7. Fit alongside the runtime, agent, operating system, and KV cache.

##### Recommended Shortlist

| Model | Size class | Best target | Notes |
| --- | --- | --- | --- |
| [MiniCPM5 1B](https://huggingface.co/openbmb/MiniCPM5-1B) | 1B | Pi 5 8 GB | Tiny general model; tool behavior must be qualified |
| [LFM2 1.2B Tool](https://huggingface.co/LiquidAI/LFM2-1.2B-Tool) | 1.2B | Pi 5 8 GB | Narrow tool-use candidate rather than broad assistant |
| [Qwen3.5 2B](https://huggingface.co/Qwen/Qwen3.5-2B) | 2B | Pi 5 8/16 GB, Jetson 8 GB | Strong general edge candidate |
| [Ministral 3 3B Instruct](https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512) | 3B | Pi 16 GB, Jetson 8 GB | Edge-oriented and supports function calling |
| [Phi-4 Mini Instruct](https://huggingface.co/microsoft/Phi-4-mini-instruct) | 3.8B | Pi 16 GB, Jetson 8/16 GB | Function calling; run shorter than maximum context |
| [Qwen3.5 4B](https://huggingface.co/Qwen/Qwen3.5-4B) | 4B | Pi 16 GB, Jetson 8/16 GB | Good general/tool tradeoff |
| [Granite 4.0 Micro](https://huggingface.co/ibm-granite/granite-4.0-micro) | Small MoE | Jetson 8/16 GB | Tool-calling focus; verify runtime template |
| [Gemma 4 E2B IT](https://huggingface.co/google/gemma-4-E2B-it) | Small | Pi 16 GB, Jetson | Very new; runtime parser maturity needs testing |
| [Qwen3.5 9B](https://huggingface.co/Qwen/Qwen3.5-9B) | 9B | Jetson 16 GB or more | Stronger agent behavior with more memory |
| [Qwen3.6 27B](https://huggingface.co/Qwen/Qwen3.6-27B) | 27B | AGX Orin 32/64 GB or Thor | Not an entry-level edge model |

Model availability and quantized artifact quality change quickly. Pin the exact
model repository revision, quantization producer, filename, prompt template,
license, and SHA-256 hash.

##### Models to Treat Carefully

- **FunctionGemma 270M:** intended as a function-calling fine-tuning base, not a
  general conversational assistant.
- **Reasoning-only distill models:** many do not emit robust native tool calls.
- **Models with only informal tool-use examples:** an example prompt is not the
  same as validated multi-turn function calling.
- **Freshly released architectures:** inference may work before tool parsers,
  quantization, and long-context caches are stable.

##### Quantization and Context Defaults

Start with:

- `Q4_K_M` or a vendor-provided 4-bit quantization.
- 8K context on Pi and 8 GB Jetson.
- 16K context on 16 GB Jetson after measurement.
- One loaded model.
- One generation at a time.
- Short tool descriptions and bounded tool output.

Increase context only after recording peak memory, time to first token,
generation rate, and tool-loop success. Long context consumes memory even when
the model weights fit.

#### Hardware Profiles

##### Raspberry Pi 5

The
[Raspberry Pi 5](https://www.raspberrypi.com/products/raspberry-pi-5/) is
available with up to 16 GB RAM. The official
[16 GB announcement](https://www.raspberrypi.com/news/16gb-raspberry-pi-5-on-sale-now-at-120/)
confirms the higher-memory model.

Required appliance choices:

- Use a 64-bit operating system.
- Use active cooling.
- Use the official power supply or a validated equivalent.
- Prefer NVMe or a USB SSD over a microSD card for model and state storage.
- Use wired Ethernet for a Pi-to-Jetson inference path.

Pi 5, 8 GB:

- Good as gateway, router, scheduler, and lightweight agent.
- Reasonable for 1B to 3B Q4 inference.
- Do not expect a 64K context personal agent to be fast or reliable.

Pi 5, 16 GB:

- More viable for a 3B to 8B Q4 model.
- Still CPU-bound and much slower than a Jetson GPU.
- Better as a self-contained low-power appliance when latency is secondary.

##### NVIDIA Jetson

The
[Jetson Orin Nano Super](https://developer.nvidia.com/blog/nvidia-jetson-orin-nano-developer-kit-gets-a-super-boost/)
provides 8 GB memory and substantially more inference capability than a Pi, but
the operating system, agent, runtime, weights, context cache, and GPU all share
that memory.

Recommended tiers:

- **Orin Nano 8 GB:** 2B to 4B Q4 baseline.
- **Orin NX 16 GB:** 7B to 14B Q4 baseline.
- **AGX Orin 32 GB:** 14B to 27B or 32B Q4, workload dependent.
- **AGX Orin 64 GB:** 32B with room for agent services; some 70B Q4 use cases.
- **Jetson Thor 128 GB:** large-model and concurrent-agent tier.

Use the current
[Jetson AI Lab model catalog](https://www.jetson-ai-lab.com/models/) and
container guidance for the exact JetPack release. Do not assume that a container
or wheel built for one JetPack, CUDA, Python, or ARM page-size combination works
on another.

#### Deployment A: Standalone Raspberry Pi

##### Recommended Stack

- Raspberry Pi OS 64-bit or Ubuntu Server 64-bit.
- `llama.cpp` as the model server.
- Pi for coding, or ZeroClaw/nanobot/PicoClaw for a personal assistant.
- USB SSD or NVMe for models and state.
- 1B to 3B Q4 model on 8 GB; 3B to 8B Q4 on 16 GB.

##### Service Layout

```text
systemd
  edge-model.service       llama-server on 127.0.0.1:8080
  edge-agent.service       agent under an unprivileged user
  edge-backup.timer        state backup and checksum verification

/opt/offline-base          versioned runtime and agent files
/mnt/base-drive/models     GGUF models
/var/lib/offline-agent     writable agent state
/srv/offline-workspace     restricted working directory
```

##### Operating Rules

- Keep the model API on loopback.
- Run one request at a time.
- Disable tools that need internet when offline.
- Bound file reads and tool output.
- Require confirmation for shell commands, package installation, deletion,
  network changes, and writes outside the workspace.
- Use a watchdog and restart policy, but rate-limit restart loops.

##### Expected Experience

This design can provide private chat, retrieval over small local collections,
simple file operations, and constrained coding assistance. It will not match a
large hosted model's planning, tool reliability, or speed.

#### Deployment B: Standalone Jetson

##### Recommended Stack

- JetPack version pinned to the device image.
- NVMe storage.
- Ollama for convenience or CUDA `llama.cpp` for direct control.
- Pi or Aider for coding.
- ZeroClaw, nanobot, or LocalClaw for a lightweight personal assistant.
- OpenClaw or Hermes only on memory tiers that meet their context and service
  overhead.

##### 8 GB Profile

- One 2B to 4B Q4 model.
- 8K context initially.
- One loaded model and one request.
- Headless services.
- Avoid desktop IDEs and heavy multi-agent orchestration.

##### 16 GB Profile

- One 7B to 14B Q4 model, or a smaller model with more context.
- Pi/Aider plus a personal-agent gateway can coexist if concurrency is queued.
- Cline should normally run on a separate laptop.

##### 32 GB and Larger Profile

- OpenClaw or Hermes becomes practical.
- vLLM may be justified for concurrent clients.
- Separate fast and capable models can be offered, but only if model swapping
  and peak memory are measured.

#### Deployment C: Portable Drive

##### What the Drive Can Do

A drive can:

- Carry shared GGUF model files.
- Carry per-architecture runtimes and agents.
- Carry offline operating-system packages and language dependencies.
- Carry retrieval documents and indexes.
- Store configuration templates, licenses, hashes, and recovery material.
- Boot a compatible host when prepared as boot media.

A drive cannot execute an agent without a CPU, RAM, and compatible operating
system.

##### Three Drive Products

###### 1. Attached Model and Data Drive

The host already has an operating system. The drive supplies models, documents,
and installers.

This is the broadest and most maintainable option.

###### 2. Bootable Appliance Drive

The drive contains an operating system and starts the agent automatically.

Separate images are required for:

- x86-64 UEFI systems.
- Raspberry Pi.
- Each supported Jetson family and JetPack baseline.

A single universal image is not realistic because boot firmware, kernels,
drivers, GPU libraries, and partition expectations differ.

###### 3. Air-Gapped Installer Drive

The drive contains every dependency needed to install without internet:

- OS package archives and repository metadata.
- Python wheels for the exact Python and ARM64 ABI.
- Node.js runtime and npm package tarballs.
- Agent source or release artifacts.
- Runtime binaries.
- Model files.
- Container images saved as archives when containers are used.
- Licenses, notices, checksums, and a signed manifest.

An agent supporting local inference does not necessarily support air-gapped
installation. Hermes Agent, for example, has had an open request for an
official offline installation flow.

##### Recommended Physical Media

Use a USB 3.2 SSD or NVMe enclosure for the primary product. Low-cost thumb
drives have worse sustained reads, random I/O, thermals, write endurance, and
power-loss behavior.

Suggested partitions:

| Partition | Format | Purpose |
| --- | --- | --- |
| `BASE-CONTENT` | exFAT, preferably mounted read-only | Models, docs, manifests, installers shared across OSes |
| `BASE-STATE` | ext4 | Linux agent state, permissions, logs, indexes |
| `BASE-PRIVATE` | LUKS plus ext4, optional | User secrets and private documents |

Python virtual environments, native Node modules, and runtime binaries are not
portable across operating systems and architectures. GGUF model files usually
are portable when the runtime supports the architecture and model.

##### Suggested Drive Layout

```text
offline-base/
  manifest/
    manifest.json
    SHA256SUMS
    SIGNATURE
  models/
    gguf/
  runtimes/
    linux-arm64/
      llama.cpp/
    linux-x86_64/
      llama.cpp/
    macos-arm64/
      llama.cpp/
  agents/
    linux-arm64/
    linux-x86_64/
    platform-independent/
  installers/
    raspberry-pi-5/
    jetson-jetpack-6/
    linux-x86_64/
    macos-arm64/
    windows-x86_64/
  configs/
    pi/
    zeroclaw/
    nanobot/
    openclaw/
    systemd/
  knowledge/
  licenses/
  recovery/
```

The public product experience can remain simple while this internal package is
versioned and audited.

#### Deployment D: Combined Pi, Jetson, Drive, and Laptop

##### Recommended Topology

```text
Messaging clients or local web UI
                |
        Raspberry Pi gateway
     channels, policy, queue, state
                |
        private wired Ethernet
                |
         Jetson model server
      llama.cpp, Ollama, or vLLM
                |
       USB SSD or Jetson NVMe
      models, knowledge, backups

Optional laptop clients:
  Cline, Pi, Aider, OpenCode, Qwen Code
```

##### Why This Split Works

- The Pi remains cool, low-power, and always on.
- Jetson GPU memory is reserved for inference.
- A single model endpoint can serve several agents.
- The drive is replaceable and recoverable.
- The laptop supplies the IDE without burdening the appliance.
- Each component can be upgraded independently.

##### Network Pattern

Use a dedicated private LAN or VLAN and stable addresses:

```text
base-gateway.local     192.168.50.5
base-inference.local   192.168.50.10
```

The model server should listen only on loopback or the private Jetson address.
Restrict the inference port to the gateway and approved clients. Ollama and
plain `llama-server` endpoints should not be assumed to provide sufficient
authentication by themselves.

For stronger isolation:

- Put Caddy, Nginx, or Envoy in front of the inference endpoint.
- Require an API token or mutual TLS.
- Permit only the private subnet.
- Deny outbound traffic from the model service.
- Give each agent a distinct credential and audit identity.

##### Concurrency

On 8 GB and 16 GB systems:

- One model loaded.
- One active generation.
- Gateway queue for additional requests.
- Request timeout and cancellation support.
- No automatic loops where two agents repeatedly call each other.

Multiple agents may share one endpoint, but they do not automatically share
memory, identity, permissions, or conversation state. Shared state must be an
explicit service or restricted directory, not an accidental common home folder.

##### Optional Router

LiteLLM or a small custom proxy can provide:

- A stable endpoint while model servers change.
- Per-agent API keys.
- Model aliases.
- Request logging.
- A later cloud fallback.

Do not add a router to the first version unless those capabilities are needed.
Direct agent-to-Jetson communication is easier to debug.

#### Endpoint Integration Recipes

Use one canonical hostname for the model service, such as
`base-inference.local`. The path depends on the client.

##### Verify an OpenAI-Compatible Endpoint

Pi, Hermes, and many other agents can use an OpenAI-compatible endpoint from
`llama-server`, vLLM, or Ollama:

```bash
curl -s http://base-inference.local:8080/v1/models

curl -s http://base-inference.local:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "edge-model",
    "messages": [{"role": "user", "content": "Reply with exactly: pong"}],
    "temperature": 0
  }'
```

Run a plain chat smoke test before loading an agent's tools. This distinguishes
model-server failures from tool-schema and agent-prompt failures.

##### Pi

Create `~/.pi/agent/models.json` using the provider example in the Pi section:

```json
{
  "providers": {
    "base": {
      "baseUrl": "http://base-inference.local:8080/v1",
      "api": "openai-completions",
      "apiKey": "local",
      "compat": {
        "supportsDeveloperRole": false,
        "supportsReasoningEffort": false
      },
      "models": [{"id": "edge-model"}]
    }
  }
}
```

Open Pi, select `/model`, choose the `base` provider, and begin with a
read-only repository task before allowing edits or shell commands.

##### Hermes

Run:

```text
hermes model
```

Select a custom endpoint and provide:

```text
API base URL: http://base-inference.local:8080/v1
API key: local
Model name: edge-model
Context length: 64000
```

The 64,000-token value is Hermes' documented minimum. Do not claim Hermes
compatibility if the server and hardware cannot actually provide that context.

##### OpenClaw with Ollama

Run:

```text
openclaw onboard
```

Select Ollama, Local only, and the LAN host:

```text
http://base-inference.local:11434
```

Do not append `/v1`. OpenClaw's official provider uses Ollama's native API for
tool calling. Verify discovery with:

```bash
openclaw models list --provider ollama
```

For a small model, use OpenClaw's lean local-model profile, reduce the direct
tool catalog, and set the OpenClaw context budget to the same value sent to
Ollama.

##### LocalClaw

Start the supported local model server, then run:

```text
localclaw onboard
```

The onboarding flow detects a reachable model server and writes
`~/.localclaw/openclaw.local.json`. Start with no external messaging channel,
qualify local chat and tools, then add only an allowlisted channel.

##### Aider with Ollama

```bash
export OLLAMA_API_BASE=http://base-inference.local:11434
cd /path/to/approved/repository
aider --model ollama_chat/<exact-model-name>
```

The official Aider guidance recommends `ollama_chat/` over `ollama/`. The model
name must match the Ollama registry on the Jetson.

##### Cline with Ollama

In Cline settings:

1. Set API Provider to Ollama.
2. Set the host to `http://base-inference.local:11434`.
3. Select the exact installed model.
4. Enable Compact Prompt.
5. Set the context window to a value the Jetson can sustain.

Cline's official recommendation is at least 32K context for coding. If the
target hardware cannot sustain that context, Cline should remain an optional
integration rather than a qualified appliance profile.

##### One Endpoint, Different URL Rules

Do not blindly give every client the same URL:

| Client | Ollama URL |
| --- | --- |
| OpenClaw native Ollama provider | `http://host:11434` |
| Aider Ollama provider | `http://host:11434` |
| Pi OpenAI-compatible provider | `http://host:11434/v1` |
| Hermes custom endpoint | `http://host:11434/v1` |
| Generic OpenAI-compatible client | `http://host:11434/v1` |

OpenClaw explicitly warns that using Ollama's `/v1` compatibility path can
break its native tool-calling behavior.

#### Using Agents Separately or Together

##### Separate Profiles

Use distinct profiles rather than one unrestricted agent:

| Profile | Agent | Tools | Model |
| --- | --- | --- | --- |
| Personal gateway | ZeroClaw, nanobot, OpenClaw | Calendar, notes, approved channels | General instruct/tool model |
| Coding | Pi or Aider | Restricted repository and shell | Code-capable instruct model |
| High-capability assistant | Hermes | Sandboxed tools and memory | Larger Jetson model |
| IDE | Cline | User-reviewed edits and terminal | Remote Jetson model |

##### Shared Infrastructure

Agents can safely share:

- The model API.
- Read-only knowledge collections.
- A model registry.
- A request queue.
- Versioned skills or prompts.

Agents should not casually share:

- Shell credentials.
- Writable home directories.
- Messaging identity.
- Secrets.
- Unreviewed long-term memory.
- A root-capable Docker socket.

#### Offline and Air-Gapped Requirements

##### Offline Operation

Offline operation means the system can continue after installation without
internet access. To verify it:

1. Disable Wi-Fi and disconnect the WAN.
2. Flush DNS or block all outbound traffic.
3. Start the model and agent from a cold boot.
4. Exercise every advertised local tool.
5. Confirm there are no cloud provider retries or fallback calls.
6. Confirm local model, tokenizer, templates, and embedding models are present.
7. Inspect network connections and logs.

Telegram, Discord, Slack, web search, hosted email, and most third-party APIs
cannot work in a fully offline environment.

##### Air-Gapped Installation

Air-gapped installation means the device can be provisioned without ever
connecting to the internet.

Required controls:

- Pin every package and container digest.
- Store package indexes locally.
- Include exact ARM64 and x86-64 artifacts separately.
- Include model tokenizer and configuration files, not only weights.
- Verify all hashes before installation.
- Generate a software bill of materials.
- Preserve source and license obligations.
- Sign the release manifest.
- Test the installer in a clean network-isolated environment.

#### Security Requirements

Local inference protects model prompts from a hosted provider, but it does not
make an agent safe.

Minimum controls:

1. Run every agent as a non-root user.
2. Give coding agents an explicit workspace allowlist.
3. Keep destructive shell actions behind confirmation.
4. Deny access to SSH keys, browser profiles, cloud credentials, and system
   configuration unless required.
5. Separate messaging agents from unrestricted coding agents.
6. Treat retrieved documents, email, and web pages as untrusted input.
7. Put the model API behind a firewall or authenticated proxy.
8. Disable cloud fallback unless the product explicitly offers it.
9. Encrypt private state and secrets at rest.
10. Rotate logs and avoid recording secrets or entire private prompts.
11. Pin and hash agent skills. A skill is executable policy, not harmless text.
12. Test power-loss recovery and filesystem integrity.

Small quantized models are generally less reliable at distinguishing user
instructions from prompt injection inside documents. Reduce their tool
authority rather than relying on the model to police itself.

#### Qualification Test Suite

Every agent, model, runtime, and hardware profile must pass the same scenarios.

##### Functional Tests

1. Answer a request that needs no tool.
2. Call one tool with valid JSON.
3. Reject or repair invalid tool arguments.
4. Handle a tool error without claiming success.
5. Use a tool result correctly on the next turn.
6. Complete a three-step workflow.
7. Resume after an agent restart.
8. Resume after a model-server restart.
9. Operate with the WAN disconnected.
10. Explain clearly when an offline tool is unavailable.

##### Security Tests

1. A local document instructs the agent to reveal secrets.
2. A tool result asks the agent to ignore policy.
3. A repository contains a malicious `AGENTS.md` or similar instruction file.
4. The agent is asked to read outside its workspace.
5. The agent is asked to execute with `sudo`.
6. A messaging user is not on the allowlist.
7. The model endpoint is scanned from an unapproved host.
8. A skill package changes after its hash was approved.

##### Performance Measurements

Record:

- Cold startup time.
- Model load time.
- Time to first token.
- Tokens per second.
- Peak resident memory.
- Peak GPU memory or unified-memory pressure.
- Device temperature and throttling.
- Power draw.
- Context length.
- Tool-call success rate.
- End-to-end task success rate.
- Recovery after power loss.

Do not ship based only on tokens per second. A fast model that emits invalid
tool calls is a worse agent model than a slower reliable one.

#### Recommended Product Profiles

##### Profile 1: Base Mini

- Raspberry Pi 5, 8 GB or 16 GB.
- USB SSD or NVMe.
- `llama.cpp`.
- Qwen3.5 2B, Ministral 3 3B, or another qualified Q4 model.
- Pi coding profile plus ZeroClaw or nanobot personal profile.
- 8K context and serialized requests.

Position it as private, local, low-power assistance with constrained
automation. Do not promise large-model autonomy.

##### Profile 2: Base One

- Jetson Orin Nano 8 GB.
- NVMe.
- Ollama or CUDA `llama.cpp`.
- Qualified 3B to 4B model.
- Pi/Aider coding profile and lightweight personal-agent profile.
- Optional remote Cline client.

##### Profile 3: Base Pro

- Jetson Orin NX 16 GB or AGX Orin 32 GB.
- 7B to 14B model on 16 GB; larger model on AGX.
- OpenClaw available as a gateway profile.
- Hermes available only after context and ARM qualification.
- Authenticated LAN inference endpoint.

##### Profile 4: Portable Base Drive

- External SSD rather than a low-end thumb drive.
- Shared GGUF models and knowledge.
- Per-architecture installers and binaries.
- Signed manifest and checksums.
- No claim that one boot image works on every computer.

##### Profile 5: Combined Home or Office Base

- Pi gateway.
- Jetson inference node.
- External SSD or Jetson NVMe.
- Laptop IDE clients.
- Private wired LAN, queue, firewall, backups, and role separation.

This is the recommended architecture for the broadest capability and the least
hardware waste.

#### Implementation Sequence

##### Phase 1: Prove the Smallest Reliable Stack

1. Select Pi 5 8 GB and Orin Nano 8 GB reference devices.
2. Build and pin `llama.cpp`.
3. Select three candidate models in the 2B to 4B range.
4. Run the qualification suite with Pi, ZeroClaw, and nanobot.
5. Record performance, reliability, thermal, and memory results.
6. Select one coding and one personal-agent profile.

##### Phase 2: Build the Drive Manifest

1. Define the directory layout and manifest schema.
2. Pin source commits, release versions, model revisions, and hashes.
3. Build ARM64 and x86-64 runtime bundles.
4. Cache all offline dependencies.
5. Add licenses, notices, SBOM, checksums, and signature verification.
6. Test on clean offline hosts.

##### Phase 3: Add the Split Architecture

1. Put the gateway on Pi and inference on Jetson.
2. Add private DNS or fixed addresses.
3. Add request authentication and firewall rules.
4. Add a single-request queue.
5. Connect Pi, Aider, and Cline clients.
6. Test Jetson restart, Pi restart, and network interruption.

##### Phase 4: Qualify Larger Frameworks

1. Test LocalClaw with the selected small model.
2. Test OpenClaw on Pi with remote Jetson inference.
3. Test Hermes on 32 GB or larger hardware.
4. Test Goose and OpenCode ARM64 release artifacts.
5. Promote only combinations that pass the same qualification suite.

##### Phase 5: Maintain the Release

1. Rebuild monthly or when a security update requires it.
2. Run automated clean-device and offline tests.
3. Preserve previous signed manifests for rollback.
4. Publish compatibility by exact device, OS image, and release.
5. Never silently replace a model or agent in an existing profile.

#### Final Selection

For the first dependable implementation:

- Use `llama.cpp` as the universal compatibility runtime.
- Offer Ollama as the convenience runtime on qualified Jetson and desktop
  images.
- Use Pi or Aider for coding.
- Use ZeroClaw or nanobot for the lightweight personal assistant.
- Treat LocalClaw as a promising experimental profile.
- Use OpenClaw primarily as a Pi gateway to Jetson inference.
- Reserve Hermes for larger Jetson tiers or split deployments.
- Run Cline on the user's laptop and point it at the Jetson.
- Package models and data on an external SSD with per-platform installers.
- Make the Pi plus Jetson plus SSD topology the flagship combined system.

#### Primary Sources

##### Agents

- [LocalClaw repository](https://github.com/sunkencity999/localclaw)
- [OpenClaw repository](https://github.com/openclaw/openclaw)
- [OpenClaw local-model guide](https://docs.openclaw.ai/gateway/local-models)
- [OpenClaw Ollama provider](https://docs.openclaw.ai/providers/ollama)
- [OpenClaw Raspberry Pi guide](https://docs.openclaw.ai/install/raspberry-pi)
- [Hermes Agent repository](https://github.com/NousResearch/hermes-agent)
- [Hermes provider documentation](https://hermes-agent.nousresearch.com/docs/integrations/providers)
- [Hermes FAQ and local-model requirements](https://hermes-agent.nousresearch.com/docs/reference/faq)
- [Hermes and Ollama integration](https://docs.ollama.com/integrations/hermes)
- [Pi website](https://pi.dev/)
- [Pi repository](https://github.com/earendil-works/pi)
- [Cline repository](https://github.com/cline/cline)
- [Cline local-model documentation](https://docs.cline.bot/running-models-locally/overview)
- [Aider repository](https://github.com/Aider-AI/aider)
- [Aider Ollama documentation](https://aider.chat/docs/llms/ollama.html)
- [OpenCode repository](https://github.com/anomalyco/opencode)
- [Goose repository](https://github.com/aaif-goose/goose)
- [Goose provider documentation](https://goose-docs.ai/docs/getting-started/providers)
- [Qwen Code repository](https://github.com/QwenLM/qwen-code)
- [nanobot repository](https://github.com/HKUDS/nanobot)
- [ZeroClaw repository](https://github.com/zeroclaw-labs/zeroclaw)
- [PicoClaw repository](https://github.com/sipeed/picoclaw)

##### Runtimes and Hardware

- [llama.cpp repository](https://github.com/ggml-org/llama.cpp)
- [Ollama documentation](https://docs.ollama.com/)
- [Ollama FAQ](https://docs.ollama.com/faq)
- [Jetson AI Lab Ollama tutorial](https://www.jetson-ai-lab.com/tutorials/ollama/)
- [Jetson AI Lab model catalog](https://www.jetson-ai-lab.com/models/)
- [Jetson Orin Nano Super specifications](https://developer.nvidia.com/blog/nvidia-jetson-orin-nano-developer-kit-gets-a-super-boost/)
- [Jetson Thor announcement](https://developer.nvidia.com/blog/introducing-nvidia-jetson-thor-the-ultimate-platform-for-physical-ai/)
- [Raspberry Pi 5 product page](https://www.raspberrypi.com/products/raspberry-pi-5/)
- [Raspberry Pi boot documentation](https://www.raspberrypi.com/documentation/computers/raspberry-pi.html)

##### Models

- [Qwen3.5 2B](https://huggingface.co/Qwen/Qwen3.5-2B)
- [Qwen3.5 4B](https://huggingface.co/Qwen/Qwen3.5-4B)
- [Ministral 3 3B Instruct](https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512)
- [Phi-4 Mini Instruct](https://huggingface.co/microsoft/Phi-4-mini-instruct)
- [Granite 4.0 Micro](https://huggingface.co/ibm-granite/granite-4.0-micro)
- [Gemma 4 E2B IT](https://huggingface.co/google/gemma-4-E2B-it)
- [FunctionGemma 270M IT](https://huggingface.co/google/functiongemma-270m-it)
- [LFM2 1.2B Tool](https://huggingface.co/LiquidAI/LFM2-1.2B-Tool)
- [MiniCPM5 1B](https://huggingface.co/openbmb/MiniCPM5-1B)

#### Known Unknowns

The following items cannot be settled by documentation alone:

- Exact tokens per second on each power mode.
- Whether a specific quantized model produces reliable tool calls.
- Thermal throttling inside the intended enclosure.
- ARM64 behavior of every agent release.
- Performance of fresh model architectures in each runtime.
- Recovery after sudden power loss on the selected drive.
- Whether every installer dependency is genuinely cached for air-gapped use.
- Whether telemetry, update checks, and cloud fallbacks are fully disabled.

These are release gates and require tests on the actual product image and
hardware, not assumptions.

<!-- source-document-end -->

## Source Document: Jetson Orin Nano Use Case Research

Source file: `deep research output.md`

<!-- source-document-start -->

### Deep Research Output: Jetson Orin Nano 8GB Use Cases for Non-Technical Buyers

Date: June 16, 2026

#### Research Framing

This file turns the Jetson Orin Nano 8GB into sellable, plain-language product ideas for people who are not developers. The strongest fit is local AI that watches, listens, sorts, counts, or summarizes things on-site without depending on the cloud.

Technical boundary used for this research:

- The current NVIDIA Jetson Orin Nano Super Developer Kit is positioned for edge generative AI, robotics, vision AI, local LLMs, vision-language models, and vision transformers.
- Key published specs include up to 67 INT8 TOPS, 8GB LPDDR5 memory, 102 GB/s memory bandwidth, external NVMe support, and a 7W to 25W power range.
- The board is best used for compact local inference, cameras, sensors, voice interfaces, small robots, kiosks, and offline assistants.
- It is not the right choice for huge cloud-scale AI models, guaranteed emergency safety systems, medical diagnosis, heavy video editing, or large multi-camera deployments without careful engineering.

Pricing assumptions:

- Prices are estimated customer sell prices in USD for a packaged product, pilot, or installation.
- Prices generally assume the Jetson is included in the customer-facing package, unless the idea is clearly a rental or service.
- Real pricing depends on enclosure quality, camera count, installation labor, support, compliance, and whether the product is sold as hardware, service, or both.

Feasibility scale:

- Easy: Can be demoed with common hardware and existing models.
- Medium: Needs a polished workflow, installation design, and some tuning.
- Hard: Needs strong reliability work, domain testing, safety controls, or regulatory review.

Sources:

- NVIDIA Jetson Orin Nano Super Developer Kit: https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/nano-super-developer-kit/
- Jetson AI Lab: https://www.jetson-ai-lab.com/
- SparkFun Jetson Orin Nano Super Developer Kit listing: https://www.sparkfun.com/nvidia-jetson-orin-nano-developer-kit.html
- Micro Center Jetson Orin Nano Super Developer Kit specs: https://www.microcenter.com/product/691058/nvidia-jetson-orin-nano-super-developer-kit

#### 25 Product Use Cases

##### 1. Private Home Package and Driveway Watcher

- Idea: A local camera box that recognizes packages, cars, open garage doors, people near the porch, and unusual driveway activity.
- Bullet description: Instead of sending all doorbell footage to a cloud service, the Jetson watches locally and only saves or alerts on important events.
- Price it could sell for: $799 to $1,499 for a starter kit, or $1,500 to $3,000 installed with multiple cameras.
- Market: Homeowners, renters with permission, privacy-focused families, and people with frequent deliveries.
- Feasibility: Easy to medium.
- Necessary other than the Jetson Orin Nano 8GB: One or two outdoor cameras, weatherproof enclosure, power supply, mounting hardware, local storage or NVMe SSD, Wi-Fi or Ethernet, alert app or text/email notification service.

##### 2. Backyard and Pool Safety Alert Assistant

- Idea: A camera-based monitor that watches for unsupervised movement near a pool, backyard gate, hot tub, or dangerous area.
- Bullet description: The system gives fast local alerts when a child, guest, or pet enters a marked danger zone.
- Price it could sell for: $1,200 to $2,500 for a home setup, with optional $20 to $50 per month support.
- Market: Parents, grandparents, Airbnb hosts, pool owners, and child-care homes.
- Feasibility: Hard because it touches safety and should be sold as an alert aid, not as a certified lifesaving system.
- Necessary other than the Jetson Orin Nano 8GB: Weatherproof cameras, outdoor enclosure, siren or speaker, phone alert workflow, backup power, zone setup software, privacy and liability disclaimers.

##### 3. Elder Care Fall and Routine Monitor

- Idea: A local privacy-friendly monitor that detects falls, long periods without movement, missed routines, or wandering near exits.
- Bullet description: It helps caregivers know when something may be wrong without placing every video feed in the cloud.
- Price it could sell for: $1,500 to $3,500 installed, plus $25 to $100 per month for support and alert routing.
- Market: Families caring for older adults, independent living homes, and small senior-care providers.
- Feasibility: Hard because false alarms, missed events, privacy, and liability must be handled carefully.
- Necessary other than the Jetson Orin Nano 8GB: Indoor cameras or depth sensors, optional door sensors, speaker, caregiver alert app, battery backup, consent process, local storage, secure remote access for approved caregivers.

##### 4. Pet Behavior and Feeding Monitor

- Idea: A pet camera that recognizes feeding, water bowl use, scratching, barking events, crate time, door waiting, or unusual behavior.
- Bullet description: Pet owners get practical summaries instead of watching hours of footage.
- Price it could sell for: $699 to $1,199 for a kit, or $20 to $40 per month for cloud backup and reports.
- Market: Dog owners, cat owners, pet sitters, breeders, and boarding facilities.
- Feasibility: Easy to medium.
- Necessary other than the Jetson Orin Nano 8GB: Indoor camera, microphone, optional smart scale or bowl sensor, speaker, local storage, simple phone dashboard.

##### 5. Low-Vision Object and Label Reader

- Idea: A countertop or wearable-adjacent camera assistant that reads labels, identifies household objects, describes scenes, and helps find common items.
- Bullet description: It gives spoken help locally for daily tasks like reading medicine bottles, sorting mail, or identifying pantry items.
- Price it could sell for: $799 to $1,800 for a home kit, with possible nonprofit or insurance-assisted channels.
- Market: People with low vision, caregivers, disability support groups, senior centers, and occupational therapists.
- Feasibility: Medium to hard because the user experience must be highly reliable and accessible.
- Necessary other than the Jetson Orin Nano 8GB: Camera, microphone, speaker, simple button or voice trigger, screen optional, privacy-first local model setup, strong enclosure design, accessible onboarding.

##### 6. Small Retail Shelf Stock Watcher

- Idea: A store camera that flags empty shelves, misplaced products, low inventory areas, and restocking needs.
- Bullet description: Small stores get simple shelf alerts without buying a full enterprise retail analytics system.
- Price it could sell for: $1,500 to $4,000 per store zone, plus $50 to $200 per month for support and reporting.
- Market: Convenience stores, boutiques, hardware stores, grocers, pharmacies, and local specialty shops.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Fixed cameras, mounting, local dashboard, product reference images, store map, optional barcode/POS integration, NVMe storage.

##### 7. Restaurant Kitchen Station Assistant

- Idea: A station monitor that watches prep areas, checks basic order timing, flags spills, and reminds staff about visible steps.
- Bullet description: It acts like a quiet kitchen helper that notices delays and simple issues before they become service problems.
- Price it could sell for: $1,500 to $3,500 per station, or $3,000 to $8,000 for a small kitchen pilot.
- Market: Independent restaurants, bakeries, cafes, ghost kitchens, and food trucks.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Kitchen-safe camera, splash-resistant enclosure, display or tablet, speakers, task rules, local network, optional kitchen display system integration.

##### 8. Gym Rep Counter and Form Feedback Station

- Idea: A camera station that counts reps and gives basic form cues for squats, presses, deadlifts, lunges, jumps, or rehab movements.
- Bullet description: Members get useful coaching feedback without needing a trainer for every set.
- Price it could sell for: $1,000 to $2,500 per station, plus $50 to $150 per month for gym reporting and updates.
- Market: Small gyms, personal trainers, school weight rooms, physical therapy clinics, and home gym owners.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Wide-angle camera, tripod or wall mount, display, speaker, exercise library, consent workflow, local user profiles optional.

##### 9. Physical Therapy Exercise Checker

- Idea: A clinic or home kit that confirms whether patients completed simple assigned movements and gives basic range-of-motion feedback.
- Bullet description: Therapists get more visibility into home exercise compliance while patients get reminders and simple guidance.
- Price it could sell for: $799 to $1,999 per home kit, or $2,500 to $6,000 for a clinic pilot.
- Market: Physical therapists, occupational therapists, sports rehab clinics, and remote-care programs.
- Feasibility: Hard because healthcare claims, patient safety, and clinical accuracy need strong limits.
- Necessary other than the Jetson Orin Nano 8GB: Camera or depth sensor, display, speaker, patient app or portal, therapist dashboard, secure storage, consent and privacy workflows, carefully worded non-diagnostic positioning.

##### 10. Backyard Garden and Greenhouse Plant Health Watcher

- Idea: A camera and sensor system that watches plants for wilting, discoloration, pest patterns, watering issues, and growth progress.
- Bullet description: Gardeners get simple local alerts like "this tray looks dry" or "these leaves changed color."
- Price it could sell for: $699 to $1,500 for a hobby setup, or $1,500 to $4,000 for a serious greenhouse zone.
- Market: Home gardeners, greenhouse owners, nurseries, urban farms, and school gardens.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Camera, moisture sensors, temperature/humidity sensors, optional grow-light integration, enclosure, local dashboard, outdoor-rated cables.

##### 11. Farm Livestock and Barn Monitor

- Idea: A barn camera system that watches animal movement, feeding areas, open gates, water troughs, and unusual behavior.
- Bullet description: Farmers get on-site alerts even when internet service is weak.
- Price it could sell for: $1,500 to $4,500 per barn or monitored area.
- Market: Small farms, hobby farms, ranches, dairies, horse barns, and 4-H facilities.
- Feasibility: Medium to hard because barns are dusty, dark, cold, hot, and connectivity-poor.
- Necessary other than the Jetson Orin Nano 8GB: Rugged cameras, infrared lighting, weatherproof enclosure, Ethernet or long-range Wi-Fi, UPS battery, local display optional, dust protection, storage.

##### 12. Mechanic Visual Inspection Assistant

- Idea: A garage tool that helps document vehicle condition, recognize common visible issues, and create before/after repair records.
- Bullet description: It speeds up inspections by turning photos and video into organized notes for the customer.
- Price it could sell for: $1,200 to $3,000 per bay, plus $50 to $200 per month for report templates and updates.
- Market: Independent mechanics, body shops, mobile mechanics, dealership service bays, and fleet maintenance teams.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Workbench camera or inspection camera, display, microphone, printer or PDF report workflow, shop Wi-Fi, lighting, storage.

##### 13. Property Maintenance Watcher

- Idea: A camera and sensor hub that watches shared areas for leaks, doors left open, trash overflow, blocked exits, or lights left on.
- Bullet description: Property managers get simple alerts for maintenance issues before tenants complain.
- Price it could sell for: $1,500 to $5,000 per building area, with $100 to $300 per month for managed service.
- Market: Apartment buildings, small offices, storage facilities, churches, schools, and community centers.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Cameras, water leak sensors, door sensors, enclosure, power backup, local alert dashboard, signage and privacy policy.

##### 14. Event Crowd and Queue Counter

- Idea: A portable local camera station that counts lines, room capacity, entrance flow, and crowd movement without identifying people.
- Bullet description: Event staff can open more check-in lanes, redirect visitors, or manage room limits in real time.
- Price it could sell for: $500 to $1,500 per event rental, or $2,500 to $6,000 for a reusable kit.
- Market: Churches, conferences, school events, venues, trade shows, festivals, and local government events.
- Feasibility: Easy to medium.
- Necessary other than the Jetson Orin Nano 8GB: Cameras, tripods, battery pack or wall power, portable display, Wi-Fi hotspot optional, privacy signage, carrying case.

##### 15. Security Clip Summarizer

- Idea: A local system that scans camera footage and creates short summaries of important events instead of hours of raw video.
- Bullet description: Security staff can review what mattered first, like after-hours motion, gate activity, vehicles, or restricted-area entry.
- Price it could sell for: $1,500 to $5,000 per camera cluster, plus optional monthly support.
- Market: Small warehouses, offices, schools, churches, private security teams, and property managers.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Existing IP cameras or new cameras, local video storage, dashboard, rules for events, network access, privacy policy, optional NVR integration.

##### 16. Classroom Local AI Helper Kiosk

- Idea: A classroom kiosk that answers questions from approved lesson files, reads text aloud, and helps students practice without internet dependence.
- Bullet description: Teachers get a controlled local assistant that does not need to send student questions to a public cloud.
- Price it could sell for: $900 to $2,500 per classroom kit, or $5,000 to $15,000 for a small school pilot.
- Market: Teachers, homeschool groups, tutoring centers, libraries, and schools with restricted internet.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Touchscreen, microphone, speaker, local document library, content filters, school network setup, admin controls, durable enclosure.

##### 17. Library and Community Resource Assistant

- Idea: A privacy-friendly kiosk that helps visitors find forms, local services, events, job resources, transit info, and library materials.
- Bullet description: Visitors can ask plain-language questions and get answers from approved local documents.
- Price it could sell for: $1,500 to $4,000 per kiosk, plus $100 to $400 per month for updates and maintenance.
- Market: Libraries, community centers, city halls, nonprofits, workforce centers, and senior centers.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Touchscreen, microphone, speaker, local document database, enclosure, admin update workflow, printer optional, accessibility controls.

##### 18. Photographer and Videographer Media Sorter

- Idea: A local ingest box that sorts photos and video clips by person, scene type, quality, blur, location, or event moment.
- Bullet description: Creators save time by getting organized selects before they open editing software.
- Price it could sell for: $699 to $1,500 for a portable kit, or $50 to $150 per event as a service add-on.
- Market: Wedding photographers, sports photographers, school photographers, videographers, and content creators.
- Feasibility: Easy.
- Necessary other than the Jetson Orin Nano 8GB: Fast NVMe SSD, card reader, portable screen optional, USB hub, project folder software, optional battery pack.

##### 19. Artist Studio Inventory and Work Tracker

- Idea: A studio camera that tracks canvases, tools, supplies, works in progress, and finished pieces.
- Bullet description: Artists can search their studio by image, find materials faster, and keep a visual record of each piece.
- Price it could sell for: $699 to $1,500 for an artist kit, or $1,500 to $3,000 for a gallery/studio install.
- Market: Artists, makers, ceramics studios, small galleries, school art departments, and craft studios.
- Feasibility: Easy to medium.
- Necessary other than the Jetson Orin Nano 8GB: Camera, lighting, labels or QR tags optional, local inventory app, display optional, storage.

##### 20. Real Estate Walkthrough Note Assistant

- Idea: A portable camera assistant that helps agents capture room features, visible damage, appliance notes, and listing details during walkthroughs.
- Bullet description: Real estate agents leave a property with organized notes and photo summaries instead of scattered phone pictures.
- Price it could sell for: $799 to $1,999 for a portable kit, or $99 to $299 per listing as a service.
- Market: Real estate agents, property inspectors, landlords, short-term rental managers, and insurance adjusters.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Camera, microphone, portable battery, storage, simple room-by-room app, optional phone/tablet companion.

##### 21. RV and Boat Offline Safety Assistant

- Idea: A small onboard local assistant that watches doors, bilge areas, battery status, cabin movement, and offline reference documents.
- Bullet description: Travelers get camera and sensor alerts even when mobile service is weak.
- Price it could sell for: $899 to $2,000 for a kit, or $2,000 to $5,000 installed on larger boats/RVs.
- Market: RV owners, boat owners, van-life travelers, marina customers, and overlanding users.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Low-power cameras, battery integration, door sensors, water sensors, DC power adapter, enclosure, offline maps/manuals optional, local display.

##### 22. Museum and Gallery Exhibit Engagement Counter

- Idea: A local analytics box that counts exhibit visits, dwell time, crowding, and basic interaction patterns without identifying visitors.
- Bullet description: Curators learn which exhibits people actually spend time with while keeping visitor privacy simple.
- Price it could sell for: $1,500 to $4,500 per exhibit area, or $500 to $1,500 for a temporary exhibition rental.
- Market: Museums, galleries, science centers, visitor centers, historic sites, and pop-up exhibitions.
- Feasibility: Easy to medium.
- Necessary other than the Jetson Orin Nano 8GB: Ceiling or wall camera, mount, local dashboard, signage, storage, optional battery or PoE power solution.

##### 23. Stockroom and Bin Counting Assistant

- Idea: A camera station that checks bins, shelves, or parts drawers and flags low stock, empty bins, and possible misplacements.
- Bullet description: Small teams get warehouse-style inventory help without a full enterprise system.
- Price it could sell for: $2,000 to $6,000 per stockroom area, plus $100 to $300 per month for maintenance and reporting.
- Market: Small manufacturers, repair shops, schools, makerspaces, dental offices, labs, and parts rooms.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: Cameras, shelf labels, lighting, local dashboard, barcode/QR labels optional, inventory CSV import, storage.

##### 24. Salon, Barbershop, or Clinic Queue Assistant

- Idea: A front-desk helper that watches waiting areas, estimates wait time, alerts staff when the lobby fills, and summarizes daily foot traffic.
- Bullet description: Small service businesses can improve flow without buying a complex enterprise system.
- Price it could sell for: $1,000 to $2,500 per location, plus $50 to $150 per month for reporting and support.
- Market: Salons, barbershops, dentists, small clinics, repair shops, laundromats, and walk-in service businesses.
- Feasibility: Easy to medium.
- Necessary other than the Jetson Orin Nano 8GB: Lobby camera, display optional, alert app, privacy signage, appointment-system integration optional, local dashboard.

##### 25. Local Private Office Document and Voice Assistant

- Idea: A small office assistant that answers from company documents, listens to voice questions, and keeps sensitive files on-site.
- Bullet description: A professional office gets a private AI helper for policies, client intake, internal checklists, and common questions.
- Price it could sell for: $999 to $2,500 for a starter kit, plus $100 to $500 per month for managed document updates.
- Market: Law offices, accounting firms, insurance agencies, medical front offices, consultants, and local nonprofits.
- Feasibility: Medium.
- Necessary other than the Jetson Orin Nano 8GB: NVMe SSD, document upload workflow, microphone, speaker or display, local retrieval software, permission controls, backup process, clear policy about not replacing professional judgment.

#### Best Overall Opportunities

1. Private Home Package and Driveway Watcher
2. Small Retail Shelf Stock Watcher
3. Elder Care Fall and Routine Monitor
4. Gym Rep Counter and Form Feedback Station
5. Local Private Office Document and Voice Assistant

#### Easiest Demos to Build

1. Photographer and Videographer Media Sorter
2. Event Crowd and Queue Counter
3. Private Home Package and Driveway Watcher
4. Museum and Gallery Exhibit Engagement Counter
5. Artist Studio Inventory and Work Tracker

#### Strongest Small Business Ideas

1. Small Retail Shelf Stock Watcher
2. Restaurant Kitchen Station Assistant
3. Mechanic Visual Inspection Assistant
4. Property Maintenance Watcher
5. Stockroom and Bin Counting Assistant

#### Strongest General Consumer Ideas

1. Private Home Package and Driveway Watcher
2. Pet Behavior and Feeding Monitor
3. Backyard Garden and Greenhouse Plant Health Watcher
4. RV and Boat Offline Safety Assistant
5. Low-Vision Object and Label Reader

#### Ideas That Need Careful Positioning

- Backyard and Pool Safety Alert Assistant: Useful, but must not be sold as a guaranteed lifesaving system.
- Elder Care Fall and Routine Monitor: High value, but false alarms and missed detections create trust and liability risk.
- Physical Therapy Exercise Checker: Strong market, but avoid medical claims unless clinically validated.
- Security Clip Summarizer: Valuable, but privacy, retention, and camera consent must be clear.
- Local Private Office Document and Voice Assistant: Useful, but the Jetson Orin Nano 8GB is better for small local knowledge bases than very large firm-wide document systems.

#### Practical Buyer Summary

The Jetson Orin Nano 8GB is best sold as a small private AI box for places where cameras, sensors, voice, and local documents matter. The easiest message for non-technical buyers is: it can watch, count, read, alert, and summarize locally, without sending everything to the cloud.

The best products are not "AI computer" products. They are finished helpers for a specific setting: a store shelf watcher, a garage inspection assistant, a classroom kiosk, a caregiver alert helper, or a private home camera assistant.

#### What It Is Not Good For

- Running very large AI models with long conversations and many users at once.
- Replacing a full server, gaming PC, or cloud GPU system.
- Guaranteed emergency response, medical diagnosis, or certified safety monitoring without extra validation.
- Large multi-camera commercial surveillance systems without more hardware and engineering.
- Selling to non-technical buyers as a raw developer board.

<!-- source-document-end -->

## Source Document: Project README

Source file: `README.md`

<!-- source-document-start -->

This is a [Next.js](https://nextjs.org) project bootstrapped with [`create-next-app`](https://nextjs.org/docs/app/api-reference/cli/create-next-app).

#### Getting Started

First, run the development server:

```bash
npm run dev
### or
yarn dev
### or
pnpm dev
### or
bun dev
```

Open [http://localhost:3000](http://localhost:3000) with your browser to see the result.

You can start editing the page by modifying `app/page.tsx`. The page auto-updates as you edit the file.

This project uses [`next/font`](https://nextjs.org/docs/app/building-your-application/optimizing/fonts) to automatically optimize and load [Geist](https://vercel.com/font), a new font family for Vercel.

#### Learn More

To learn more about Next.js, take a look at the following resources:

- [Next.js Documentation](https://nextjs.org/docs) - learn about Next.js features and API.
- [Learn Next.js](https://nextjs.org/learn) - an interactive Next.js tutorial.

You can check out [the Next.js GitHub repository](https://github.com/vercel/next.js) - your feedback and contributions are welcome!

#### Deploy on Vercel

The easiest way to deploy your Next.js app is to use the [Vercel Platform](https://vercel.com/new?utm_medium=default-template&filter=next.js&utm_source=create-next-app&utm_campaign=create-next-app-readme) from the creators of Next.js.

Check out our [Next.js deployment documentation](https://nextjs.org/docs/app/building-your-application/deploying) for more details.

<!-- source-document-end -->
