Philip Bowers
Philip Bowers · Temple, Texas

Built from the service drive

Fifteen years running dealership fixed operations — service, parts, warranty, full P&L. Then fourteen months building the software that department actually needed. I am the domain expert and the engineer in one hire.

15 yrsFixed operations director across domestic, import and luxury
$2.7MManufacturer warranty audit closed at a $7,000 penalty
3Platforms designed, built and shipped without a team
10Repositories under CI, secret scanning and branch protection
01The gap

The people who understand the service drive cannot build. The people who build have never stood in one.

That gap is why so much dealership software gets adopted in week one and abandoned in week three. It is designed by people who have never had a technician quit mid-shift, never defended a manufacturer audit, never watched an advisor quietly go back to a clipboard because the tool added ninety seconds to a write-up.

I have done all three. I have also been on the other side of the table as the buyer — I have negotiated, installed and driven adoption of myKaarma, DealerFX, wiADVISOR and Rapid Recon end to end — in the stores I ran, and at group level across dozens of rooftops. I know exactly where these products lose people, because I have been the one who had to make them stick.

Then I started building. Not prototypes — a multi-tenant platform with release governance, an AI system with a written evaluation charter, and reporting that is in daily use across a dealer group right now.

02What I have built

Praxis

A personal AI operating platform — server, web and Wear OS, self-hosted

Personal system · not a product

Praxis is mine. It is not a product, it is not for sale, and it is not running in anyone’s dealership. It is the system I run my own day on, and it is where I work out how an AI agent can be given real authority safely — long before that pattern goes anywhere near a business. Mail, calendar, tasks and knowledge come in, it reasons about what actually matters, and inside authority I granted in advance, it acts. It runs entirely on hardware I own — a purpose-built engineering workstation — reachable only over a private tailnet.

It is also, by a clear margin, the largest thing I have built — roughly two and a half times the codebase of Garage.

Scale
184 commits · 1,052 files under version control · ~155,000 lines across a Fastify server, a React web app and a Kotlin Wear OS companion · 49 SQL migrations · 104 design and review documents

The Engineer execution plane

Praxis could already reason about software work. It could not do any of it — every Agent Room specialist runs with no filesystem, no shell and no git, deliberately. The Engineer plane is the governed layer that lets me say, from my phone:

Diagnose the problem, implement the correction, run all the tests, have the reviewer look at it, fix anything material, commit it, and deploy it.

— and have real work happen on the workstation while I am standing on a service drive.

Phone one instruction anywhere tailnet Praxis server CONTROL PLANE runs · stages · events project registry immutable authorization job queue + leases outbound poll only nothing new listens on a port Worker THE HANDS one job, leased isolated git worktree test · build · deploy streams stage events spawn Builder → Reviewer guard hook = the real boundary The worker never accepts an inbound connection. The honest cost is that the workstation has to be powered on and on the tailnet — there is no Wake-on-LAN and no cloud fallback, and the interface says so rather than implying otherwise.
  • The boundary is enforced, not requestedPrompt text asking an agent to stay in a directory is not isolation. Every agent process is spawned with settings sources disabled — so a repo-controlled settings file, which is attacker-controlled in the threat model, is never read — plus MCP servers off, slash commands off, an explicit tool allowlist, and a PreToolUse guard hook as the actual capability boundary.
  • Verified, not assumedThat isolation behavior was tested empirically against the installed CLI before the design was built on top of it, rather than taken on faith from documentation.
  • Every run is isolated and durableA fresh git worktree per run, one job claimed at a time under a heartbeated lease, an immutable mission-authorization snapshot, and stage events streamed back into durable tables — so a run survives a restart and can be replayed.
  • No shellProcesses are spawned with an argv array and a sanitized environment. There is no shell string to inject into.

Agent Room

Six named specialists — a chief of staff who breaks down an objective and picks the team, plus calendar, inbox and communications, research and two more — addressable in five modes: Direct, Delegate, Compare, Review and Council. Work moves through visible states, survives a restart, and can be cancelled or safely retried. No specialist touches the outside world directly; every external action routes through the approval engine below.

Program 5C — the safe action engine

This is the authority boundary between a recommendation and a real mutation, and it is the piece of my work I would put in front of the hardest reviewer you have.

praxis.authority.v1 — the deterministic policy an LLM cannot touch
TierRuleExamples
READ_ONLYNo action request needed. Authenticated query APIs stay read-only.Context and search queries
INTERNAL_LOW_RISKA deterministic policy approval is stored. No human needed unless the action promotes or suppresses durable knowledge.Complete or defer a task; handle, dismiss or snooze an intelligence item
EXTERNAL_REVERSIBLEExact human approval required, bound to the snapshot hash.Create a Gmail draft or draft reply
EXTERNAL_HIGH_IMPACTExact human approval required, and an ambiguous outcome is never blindly retried.Send mail; update, decline or cancel a calendar event
PROHIBITEDRejected before an action request is even created. Not a guardrail — an absence of capability.Money transfer, bill payment, destructive account action
proposedprepared shown approved claimedexecuted verified HUMAN · BOUND TO THE SNAPSHOT HASH once, under a lease edited → invalidated rejected · expired failed → retried append-only event journal every transition above is written here — and database triggers reject UPDATE and DELETE on snapshots, approvals and events Sensitive content stays inside the authenticated snapshot. Audit rows carry bounded metadata, never a copy of the email.
  • Approval is bound to exact contentAction snapshots are append-only and carry a canonical SHA-256 hash. An approval is bound to one snapshot ID and that hash — so an action a human approved cannot quietly become a different action before it executes.
  • The model cannot grant itself authorityAn LLM cannot submit or override the authority tier, whether approval is required, the policy rule, the execution capability, or who approved. The caller supplies an action type; a versioned deterministic policy decides everything else.
  • Some things are simply unavailableA prohibited tier — money transfer, bill payment, destructive account actions — is rejected before an action request is even created. Money never moves autonomously, by construction rather than by prompt.
  • Exactly once, or visibly not at allDurable idempotency keys carry request hashes; reusing a key with different content is rejected. Execution is claimed under a lease, and an ambiguous outcome on a high-impact action is never blindly retried.
  • The audit trail cannot be editedDatabase triggers reject updates and deletes on snapshots, approvals and events. Sensitive content stays inside the authenticated snapshot rather than being copied into audit rows.

Knowing what not to build

The notification engine is specified in full — durable-first creation, no silent drops, at-least-once delivery with deduplication, loud failure, a dead-man's switch for the case where Praxis itself is the only thing still notifying me — and then marked, at the top of the document, requirements only, not authorized for implementation, because the security roadmap outranks it.

Writing a design that good and then not starting it is a harder decision than building it, and it is the one I would point to if you asked whether I can be trusted with a roadmap.

Stack
TypeScript · Fastify · Drizzle · SQLite · React · Kotlin (Wear OS) · Tailscale Serve · WebAuthn · squash-only trunk with linear history · GitHub Actions pinned to commit SHAs · CODEOWNERS · Dependabot

Garage

A multi-tenant accessory selling platform for dealer groups

Shipped · in use

Accessory sales are the highest-margin dollars on a new car deal and almost nobody captures them, because the process depends on a salesperson remembering a catalog. Garage puts a tablet in the customer's hands: pick the vehicle, see only what actually fits it, see the complete installed price, and hand a build to the sales desk before the deal is signed.

Every dealership runs its own branded storefront, its own pricing and its own availability — on top of one shared catalog of vehicle and fitment facts the platform owns.

The Garage accessory catalog for a 2026 Jeep Wrangler, showing fitment-filtered accessories with complete installed pricing and install timing.
Catalog filtered to one configured vehicle. Every price is the complete installed price — parts and labor, no line items to explain.
Vehicle selection step showing model choices.
Guided vehicle selection drives fitment.
Accessory detail view with installation timing and pricing.
Install timing is shown before the customer commits.
Build summary handed to the sales team, showing the vehicle, accessories, install plan and installed total.
The handoff. What today's install is, what needs an appointment, and one installed total the desk can work with.
  • The browser never picks the tenantDealership context is resolved server-side from an active membership. A Master Owner has no implicit default store and must select one explicitly — that selection is audited and reauthorized on every request. No customer workflow accepts a dealership ID from the client.
  • Immutable catalog releasesPublishing writes a new immutable release and advances a single current-release pointer inside one transaction. Rollback moves the pointer back. A failed import cannot change what is live, and a monotonic revision rejects stale publish previews.
  • Shared facts, store-owned decisionsThe platform owns vehicles, accessories and fitment. Each store owns its sell/withdraw call, installed price, availability, stock, lead time and timing — overlaid on stable accessory IDs so a store's commercial decisions survive the next platform release untouched.
  • Money is decided on the serverQuotes and handoffs are recalculated from the current release and the dealership's own policy. Client-supplied pricing is ignored. A projection layer keeps internal financials out of customer-facing payloads, enforced by sentinel-value tests.
  • Offline that refuses to leakThe service worker caches tenant-neutral static assets and a neutral offline page only. It will not cache authenticated, tenant-specific, tokenized or customer-document HTML. An in-progress build stays in local storage.
Stack
Next.js 15 App Router · TypeScript (strict) · Prisma · Better Auth + MFA · Zod · ExcelJS · Radix · PWA · Vitest · Playwright · Cloudflare · GitHub Actions

CLAIM

Warranty audit exposure and supported-recovery intelligence

Built · demonstrable today

A manufacturer warranty audit can take seven figures out of a store, and the exposure is created months earlier on repair orders nobody re-reads. CLAIM reads them, grounds every finding in the manufacturer's own published authority, and flags the exposure while it is still fixable.

I built the grounding corpus from the Stellantis rulebook: the Warranty Administration Manual, NVP chargeback review and appeal worksheets, published labor times, GCS message codes, the WCC list, the dealer policy manual, the goodwill grid and the pre-authorization bulletins. That corpus is the moat. It is also the reason a finding can be defended in front of an auditor instead of merely asserted.

  • A falsifiable charter, in writingThe pilot is governed by a written charter with a thesis stated so it can be disproved, a controlled false-positive rate and a measured — not assumed — false-negative rate. A clean negative result is defined up front as a valid outcome.
  • The system recommends, a person decidesNo autonomous claim submission. No autonomous OEM policy adjudication. Human authority at every external action, with the decision, the evidence and the source authority attributable after the fact.
  • Egress is a policy, not a defaultOne permitted egress path: sanitized, minimum-necessary, pseudonymized review payloads to a provider explicitly approved for that dealership, with provider, model, timestamp, payload classification and human authorization all recorded. No raw dealership or customer data in version control, ever.
  • Evaluated, not demoedA/B test sets with a sealed answer key, a shadow-pilot harness that runs alongside current practice without touching it, and a coaching index built from real technician stories.
  • Language disciplineCandidate opportunity is never reported as realized revenue and audit protection is never counted as recovered revenue. The product is called audit protection and supported recovery — never "maximization," never "found money."
Approach
Retrieval over a curated authority corpus · Python · Streamlit evaluation surface · shadow-pilot harness · sealed-key A/B sets · provenance and confidence on every claim

Multi-dealership data platform

Fortellis-connected DMS data, from one rooftop to thousands

In design

Group-level fixed ops reporting is usually a person exporting spreadsheets on the first of the month. What a group actually needs is every store’s numbers in one place, where a leak shows up as a figure instead of an argument. The blocker is getting DMS data out reliably. Rather than scraping nightly file drops, this design pulls CDK Drive through Fortellis — CDK's own API network — using the dealer's credentials, so the integration is sanctioned, versioned and survives a DMS upgrade. Because a store is just another credentialed source, the same architecture serves a single rooftop or several thousand; nothing about it assumes a group size.

From there it is a conventional warehouse, and the value is entirely in the modeling decisions.

CDK Drive 1 → thousands dealer-held credentials Fortellis sanctioned API versioned no file scraping Raw zone never edited never deleted rebuild from here dbt staging intermediate marts Power BI reads marts only one truth SOURCE IMMUTABLE MODELED CONSUMED Long term: a dealer data lake — service, parts, sales, F&I and OEM program data in one place, so a group can finally ask a question that crosses two departments and get one answer.
  • Absorption from the ledgerAbsorption is computed from the general ledger, not summed off repair orders. Almost every dealer dashboard gets this wrong and the number quietly disagrees with the financial statement.
  • Split repair ordersOne RO carrying customer-pay, warranty and internal lines is the normal case, not the exception. It is modeled that way from the start rather than patched later.
  • Pay type is a lookup, never a hardcodePay-type crosswalks live in seed tables under version control, and lookups that change meaning over time are modeled as slowly-changing dimensions so last year's report still reconciles.
  • Three words people mix upTechnician proficiency, efficiency and productivity are defined separately and calculated separately. Conflating them is how a shop convinces itself it is healthy.
  • Safeguards Rule from day oneDMS credentials and encryption keys live in a secrets manager, never in code or on a laptop. Access, retention and monitoring are designed against the FTC Safeguards Rule rather than retrofitted after an assessment.
Stack
Fortellis / CDK Drive APIs · Python · object storage raw zone · BigQuery or Snowflake · dbt Core · Power BI · secrets manager · dbt tests + freshness monitoring

Reporting in daily use

Built for one store, adopted across the group

In production use

Fixed Ops Tracker. Gross, RO count, customer-pay/warranty/internal mix, ELR, advisor performance and trend in one view a manager can run a meeting from. Built for my department; now used across multiple stores in the group.

CSI disposition report. Generates overnight so every advisor walks in to their own survey dispositions instead of hearing about a bad score weeks later, with a workflow that has them work the list and turn it back in weekly. The store had failed CSI two years running before it existed and has not failed a month since.

03How I work

Solo does not mean unserious.

Everything is governed

Ten repositories under a private organization, each with secret scanning, push protection, protected branches and CI. Work merges through pull requests, including my own.

Releases are identifiable

Production releases are immutable and tagged to the exact deployed commit, read from the machine's own release record rather than inferred — so a tag cannot disagree with what is actually running.

Incidents get written up honestly

When production went down I wrote a review that disputed its own premises, named the design defect rather than defending the safeguard, and escalated the business call — restore service or preserve release identity — to the owner instead of guessing.

Cost discipline on models

Agent work is routed by cost tier: the cheapest model for reconnaissance, mid-tier for implementation, the expensive one reserved for the final compliance gate. Escalation has to be justified in writing.

Tested where it matters

Unit tests on the logic, Playwright on the flows, and migration verification scripts that prove a change is additive on a disposable copy before it touches anything real.

It goes down to the metal

I build and tune the machines this runs on, and the hardware habit came first — custom workstations, thermal and power tuning, and the electronics underneath: Raspberry Pi, microcontrollers, sensors. When a problem turns out to be electrical rather than logical, that is not a handoff.

Drones, from parts to flight

Built from components and programmed from scratch — airframe, flight controller, tuning, and the control software on top. It is where I learned that the system you assembled yourself is the one you can actually debug.

Fact and forecast stay separate

Every system I build stamps a value with its source, timestamp and confidence. A prediction is never allowed to render as a known fact — particularly where money is involved.

04The domain record

Why my product judgment is worth something.

$2.7M → $7KManufacturer warranty audit at my store, closed at a $7,000 penalty. I rebuilt RO story standards, POP review, documentation discipline and warranty clerk training to get there.
2 yrs → 0The store had failed CSI two years running when I arrived. It passed the month I got there and has not failed once since.
35% → 71%First-year owner retention. New vehicle buyers coming back to us for their service work instead of going somewhere else.
+30% YoYParts and labor growth, with the highest line 60 selling gross of any store in the company.
DozensRooftops running myKaarma, DealerFX, wiADVISOR and Rapid Recon that I onboarded — negotiating the agreement, running the rollout, driving adoption. I have been the buyer these products are sold to.
4 directorsFour fixed operations directors reported to me at Lithia, across multiple stores and multiple brands on the West Coast. At my current group I support several more, and assist sister stores through manufacturer audits.
CDKExpert level, including custom build reports and dashboards written from scratch. Also Reynolds & Reynolds, DealerConnect and WiTECH. Certified with Stellantis, and previously Honda, Acura, Mercedes-Benz, BMW and GM.
05If you hired me

What I would do in the first ninety days.

Not a philosophy. This is what I would actually spend the time on, and none of it requires a ramp-up period where somebody explains the service department to me.

  1. Sit in three service drives, not three meetings.

    Different brands, different volumes, different DMS. Watch where your product gets abandoned in week three and why nobody filed a ticket about it. I can do that without a translator and without burning a customer's goodwill on a discovery call.

  2. Read the support queue as a roadmap.

    The roadmap is already written in there. What it needs is somebody who can tell the difference between a real defect, a training gap, and a manager who is angry about something else entirely — because I have been all three of those customers.

  3. Reconcile every number the product reports against the financial statement.

    Any metric that quietly disagrees with the dealer's own statement will cost a renewal eventually, and absorption is the one that is wrong most often. That audit is cheap now and expensive later.

  4. Ship something a service manager notices inside the first month.

    Small, real, and visible on the drive. Credibility with a dealership is not earned in a QBR deck; it is earned the first time something you changed saves an advisor ten minutes on a Monday.