We ship production AI you can actually depend on.

bold.black is a consulting studio for production AI. We design and ship agentic software, automate the workflows that drain your team, and embed senior engineers when you need to move now.

/ AI software development / Agentic engineering / Workflow automation / Staff augmentation
Services

Four ways we move your roadmap forward

We work across the full lifecycle — strategy, build, and the boring-but-critical production engineering that makes AI actually reliable.

01

AI software development

End-to-end product engineering with LLMs at the core — from RAG pipelines and evals to the UI your users actually touch.

  • RAG, retrieval & vector search
  • Model evaluation & guardrails
  • Full-stack product delivery
02

Agentic engineering

Autonomous and human-in-the-loop agents that plan, call tools, and complete real work — built to be observable, testable, and safe.

  • Multi-step tool-using agents
  • Orchestration & state management
  • Tracing, evals & observability
03

Workflow automation

We find the repetitive, high-cost workflows hiding in your operations and replace them with reliable, monitored automation.

  • Process discovery & mapping
  • Integrations across your stack
  • Monitoring & failure handling
04

Staff augmentation

Senior AI and platform engineers who plug into your team, ship in your codebase, and level up your people while they're there.

  • Embedded senior engineers
  • Architecture & code review
  • Knowledge transfer built in
Case study

Piranesi

internal platform

A self-contained agentic development environment

Piranesi is a sovereign platform where AI agents do real engineering work — write code, run CI, and ship to Kubernetes — entirely on infrastructure you own, with zero public endpoints.

We built Piranesi to answer the question every serious team is now asking: how do you put autonomous AI agents to work without handing your code, secrets, and infrastructure to someone else's service? Piranesi packages an entire software organization — source control, CI/CD, chat, workflow orchestration, and the agents themselves — into a single reproducible EKS cluster that bootstraps from one CloudFormation template and a set of agent skills. Everything runs behind a private network. Because security and data isolation are architectural defaults — not bolt-ons — the same design maps cleanly onto HIPAA, SOC 2, and similar regulatory frameworks.

01

Sovereign by design

Code, models, and data stay on infrastructure you own and can audit. After hand-off, the platform even hosts its own source of truth — no external dependency on day two.

02

Agents that actually ship

Hermes agents don't just chat — they read the repo, push commits, trigger CI, and deploy through GitOps. The full engineering loop, running autonomously.

03

Reproducible & disposable

Stand the whole environment up — or tear it down — from declarative infrastructure and gated agent skills. No snowflake servers, no manual runbooks.

04

Compliant by construction

Private networking, full audit trails, and infrastructure you own and control are built in from day one — so Piranesi is designed to satisfy HIPAA, SOC 2, and similar audits by its nature, not retrofitted to pass them.

// stack EKS · GitOps · agents
  • EKS Auto Mode
    self-provisioning Kubernetes, no node ops
  • ArgoCD GitOps
    declarative, self-healing workloads
  • Tailscale
    private, encrypted access — no public endpoints
  • Forgejo + Actions
    self-hosted Git & CI with container builds
  • Hermes on Bedrock
    AI agents that operate over Matrix chat
  • Temporal
    durable workflow orchestration
Open source

Easy containerized agent environments

Harness runs coding agents inside a hardened, sandboxed container — point one at a project and let it work without handing over your whole machine.

We open-sourced the sandbox we wanted for ourselves. Harness wraps Docker around three open-source coding agents — pi, opencode, and hermes — behind a single CLI. Every run is capability-dropped, signature-verified, and supply-chain hardened, so you get the upside of autonomous agents without the blast radius. It runs locally against LM Studio by default, or any major cloud model with one flag.

01

Agents, contained

Each run is locked to a single mounted directory in a capability-dropped container with no-new-privileges and a seccomp profile — the agent never touches the rest of your machine.

02

Trust the supply chain

Images are signed and verified with cosign + SLSA provenance on every run, and dependencies sit out a 7-day cooldown to dodge freshly-published supply-chain attacks.

03

Your models, your call

Local-first with LM Studio out of the box, or drop in a key for Anthropic, OpenAI, OpenRouter, Gemini and more — same CLI, same flow either way.

// features MIT · npx · docker
  • Sandboxed by default
    cap-drop ALL, no-new-privileges, seccomp
  • Three agents, one CLI
    switch pi · opencode · hermes with -a
  • Supply-chain hardened
    cosign + SLSA provenance, 7-day dep cooldown
  • Local-first
    LM Studio by default, any cloud model with -e
  • Stateful or one-shot
    persist sessions or run fully ephemeral
  • Zero install
    npx @boldblackai/harness just works
try it
$ npx @boldblackai/harness -p "write a fizzbuzz in Go"
Open source

Long-running coding agents that live in Slack

bclaw spins up your own long-running "claw" — an opinionated hermes-agent deployment — inside your Slack workspace, running entirely on an AWS account you own.

bclaw — short for "BusinessClaw" — is a CLI that scaffolds a complete repository for one long-running agent, each with its own Slack app and identity. Generate a @swe-pal for code reviews and pull requests, or a @reportclaw that posts scheduled reports to the channels you choose — create, customize, and deploy as many as you like. Every claw runs hermes-agent on our hardened harness image via ECS Fargate, with Slack and GitHub integration built in, state persisted and backed up on EFS, and everything deployed into infrastructure you control.

01

A claw in your workspace

Each bclaw is one long-running hermes-agent with its own Slack app and identity — spin up a @swe-pal for code reviews and PRs, or a @reportclaw for scheduled updates, right where your team already works.

02

Production-shaped from day one

Every claw ships on the same hardened foundation we run ourselves: hermes-agent on ECS Fargate, Slack + GitHub wired in, and state persisted and backed up on EFS — no glue code, no half-finished infrastructure.

03

Your account, your control

Everything deploys into your own AWS account — a CloudFormation stack, scoped IAM roles, KMS-encrypted SSM, and an isolated ECS cluster per claw — generated, managed, and torn down with gated skills.

// features MIT · npx · aws
  • Lives in your Slack
    an opinionated hermes-agent claw in your workspace
  • hermes-agent on Fargate
    runs on our hardened harness image via ECS
  • Slack + GitHub built in
    reviews code and opens PRs where your team works
  • Persisted & backed up
    state on AWS EFS survives restarts and updates
  • Your account, your control
    CloudFormation, scoped IAM, KMS-encrypted SSM
  • Skill-driven lifecycle
    scaffold, /setup, /manage, /teardown — one claw per repo
try it
$ npx @boldblackai/create-bclaw swe-pal
Selected work

Teams that trusted us to ship

From stealth startups to established enterprises, we partner with teams who treat AI as core infrastructure — not a demo.

agent platform Stealth Fintech
RAG copilot Series B SaaS
ops automation Global Logistics Co.
intake agents Healthcare Network
eval harness DevTools Startup
diligence agents Private Equity Firm
20+
AI systems shipped to production
10x
faster cycle times on automated workflows
100%
senior engineers — no juniors billed
How we work

A short path from idea to production

No bloated discovery phases. We de-risk fast, ship the real thing, and leave you with something your team can run.

  1. 01

    Scope

    A focused working session to pin down the problem, the constraints, and what 'done' looks like. You leave with a concrete plan, not a sales deck.

  2. 02

    Prototype

    We build the riskiest slice first — a working prototype against your real data — so we prove value before anyone commits to the full build.

  3. 03

    Ship

    Production engineering with evals, observability, and CI baked in. We deploy into your stack and hand over something your team can own.

  4. 04

    Scale

    Monitoring, iteration, and knowledge transfer. We make ourselves replaceable — or stay on as the embedded team you can scale with.

Start a project

Tell us what you're building

Drop the details below. We read every message and reply within one business day — usually with a few sharp questions and a sense of how we'd approach it.

Senior engineers only — you talk to the people building.
Fixed-scope projects or embedded engagements.
We'll tell you if we're not the right fit.