A privacy-preserving multi-agent framework, built on LangGraph, that separates cloud-based reasoning from edge-based execution through a two-layer Sanitization Bridge.
MACE lets an organisation use a commercial cloud LLM (e.g., Gemini) for what it is best at (intent analysis, multi-step planning, checkpoint evaluation, response-template authoring) while all tools, data operations, and response generation run on local infrastructure with small language models (Ollama or any OpenAI-compatible endpoint). The two sides communicate exclusively through operation commands (cloud to edge) and fuzzy status descriptions (edge to cloud). Raw data never crosses the bridge.
| Feature | Description |
|---|---|
| Two-layer Sanitization Bridge | Layer 1: rule-based whitelist reconstruction of structured results via fuzzification templates (no LLM, injection-immune, fail-closed). Layer 2: every nonempty text uses local-model semantic sanitization, regardless of length; an unavailable model, failed call, or empty output raises SanitizationError and aborts the workflow. No rule-only fallback is used. |
| Cloud-edge hybrid LLM | Orchestrator runs on a cloud LLM; edge agents and the Safety Agent run on local SLMs. |
| Dynamic agent registry | Agents register capability descriptions (and optional per-agent serving endpoints) at runtime; the Orchestrator discovers them from the registry, never from hard-coded lists. |
| Plan-embedded checkpoints | The Orchestrator inserts checkpoint steps into its plans; at each checkpoint it evaluates sanitized progress summaries and decides continue / stop / modify (with replacement steps). |
| Chained peer re-delegation | An edge agent that cannot finish a task delegates it to a peer; the executor guards against self-delegation and routes duplicate delegations through an adjudication checkpoint. Peers may delegate further. |
| Fault tolerance | Transient errors (HTTP 408/429/5xx) retry in place; parameter errors route back to the Orchestrator for bounded re-planning with the failure reason attached. |
| Per-agent LLM config with fallback | Each agent has a default provider and an optional fallback (gemini / ollama / openai_compatible), configurable per agent via environment variables. The Safety Agent never uses provider fallback and rejects known cloud providers; other roles do not restrict fallback targets; no fallback is applied unless one is configured. EDGE_LOCAL_ONLY restricts every role except the orchestrator to local providers. |
| Identity token threading | The end user's identity token is threaded through state to edge agents, which attach it to internal API calls (Authorization: Bearer ...). The framework holds no credential of its own and the token never enters a cloud prompt. Provider-agnostic (no identity-provider dependency). |
| Secure logging | Optional log masking of sensitive fields and patterns (tokens, emails, card numbers). |
+------------------------------------------------------------------+
| User / Frontend |
| (message + end-user identity token) |
+---------------------------------+--------------------------------+
|
v
+------------------------------------------------------------------+
| Cloud Strategy Layer (Orchestrator) |
| |
| - Intent analysis, multi-step planning Cloud LLM (Gemini / |
| - Checkpoint evaluation (continue/ OpenAI-compatible) |
| stop/modify) |
| - Response-template authoring NEVER sees raw data |
| NEVER sees the token |
+---------------------------------+--------------------------------+
commands, templates | ^ fuzzy status descriptions
v |
+-----------------------------------+
| Sanitization Bridge |
| L1: whitelist reconstruction |
| (templates, no LLM, |
| fail-closed) |
| L2: Safety Agent (local SLM |
| semantic compression + |
| pattern filtering) |
+-----------------------------------+
| ^
v | raw results
+------------------------------------------------------------------+
| Edge Execution Layer (trusted side) |
| |
| Step Executor: retries, re-planning, dynamic task insertion, |
| checkpoint routing, context injection |
| |
| +-----------+ +-----------+ +-----------+ |
| | Agent A | | Agent B | | Agent C | ... |
| | (queries) | | (reports) | | (calc) | |
| +-----------+ +-----------+ +-----------+ |
| |
| - Full access to private data - Local SLM (Ollama / |
| - Internal API calls made with OpenAI-compatible) |
| the user's own token - Final responses generated |
| on the edge |
+------------------------------------------------------------------+
| Cloud (Orchestrator) | Edge (agents) |
|---|---|
| Task planning | Access private data |
| Intent analysis | Execute tools and API calls |
| Checkpoint evaluation | Run calculations |
| Template authoring | Generate detailed responses |
| No data access, no tool execution, no user token | Holds the full reality |
git clone https://github.com/watsonshih/MACE.git
cd MACE
pip install -e . # core only; dry-runs offline
python examples/quickstart.pyThe quickstart runs a full pipeline of plan, edge execution, sanitization
(raw vs. sanitized shown side by side), checkpoint decision, and final response,
against an in-memory toy inventory. No cloud key is required: without
GEMINI_API_KEY (or with MACE_OFFLINE=1) a deterministic offline planner
stands in for the cloud LLM.
To use the real cloud Orchestrator and a local edge model:
pip install -e ".[gemini,ollama]"
cp .env.example .env # fill in GEMINI_API_KEY
ollama pull qwen3.5-9b
python examples/quickstart.pyOr with Docker:
docker compose up --buildimport asyncio
from mace import (
MACEService, BaseEdgeAgent, AgentCapability, AgentModelConfig,
AgentState, SanitizationBridge, bearer_headers,
)
class MyDataAgent(BaseEdgeAgent):
async def process(self, state: AgentState) -> AgentState:
task = self.get_task_description(state)
headers = bearer_headers(state) # user's token, pass-through
data = await my_internal_api.search(task, headers=headers)
self.set_result(state, success=True, data=data,
message=f"Found {len(data)} records",
action_type="list_records")
return state
bridge = SanitizationBridge(sensitive_fields={"salary", "revenue"})
service = MACEService(bridge=bridge)
service.register_agent(
MyDataAgent(name="data_agent"),
AgentCapability(
name="data_agent",
display_name="Data Agent",
description="Queries the local database for records.",
capabilities=["search_records", "get_record_by_id"],
model_config=AgentModelConfig(default_model="ollama"),
),
)
async def main():
print(await service.chat("Find all records from last month",
user_token=end_user_token))
asyncio.run(main())Raw edge result (stays on the edge):
{
"success": true,
"status": "success",
"data": [
{"item": "Widget B", "stock": 8, "unit_cost": 47.0,
"supplier_email": "orders@supplier-b.example"},
{"item": "Gadget C", "stock": 3, "unit_cost": 230.0,
"supplier_email": "contact@supplier-c.example"}
],
"message": "Found 2 items below their reorder point",
"action_type": "stock_query"
}What the cloud Orchestrator receives instead:
{
"status": "success",
"fuzzy_description": "Stock query returned 2 items",
"data_available": true,
"record_count": 2,
"has_error": false
}Layer 1 does not filter the raw payload; it reconstructs a brand-new
object from a fixed schema and a whitelisted template. The only values derived
from the raw result are a status flag, an integer count, an optional numeric
rate, and (on failure) an error category from a fixed enumeration. Prompt
injection payloads embedded in tool output therefore cannot survive the bridge,
and an unrecognised action_type fails closed to a status-only summary
("Operation completed").
The default template set contains fourteen action keys
(list_records, get_record, record_created, list_related_records,
list_entries, entry_created, search_result, knowledge_query,
calculation_completed, report_generated, completeness_check,
external_fetch, error, default), mirroring the fourteen registered by
the production deployment one-for-one. Production keys were domain-specific
(several were per-resource instances of the same list/create archetypes); the
defaults here are their domain-neutral equivalents, and applications are
expected to register their own per-action templates via fuzzy_templates.
All settings come from environment variables or a .env file
(see .env.example).
| Variable | Default | Description |
|---|---|---|
GEMINI_API_KEY |
(none) | Cloud LLM key (planning). Optional: without it, the orchestrator uses its configured fallback, if any. |
GEMINI_MODEL |
gemini-3-flash-preview |
Cloud model name. |
GEMINI_TIMEOUT / GEMINI_MAX_RETRIES |
20 / 5 |
Per-call timeout (s) and SDK retries. |
OLLAMA_BASE_URL |
http://localhost:11434 |
Edge LLM server. |
OLLAMA_MODEL |
qwen3.5-9b |
Edge model name. |
OLLAMA_NUM_CTX |
128000 |
Edge context window (max tokens). |
OPENAI_COMPATIBLE_BASE_URL / _API_KEY / _MODEL |
(none) | Any OpenAI-compatible endpoint (vLLM, llama.cpp, MLX serving, ...). |
AGENT_{NAME}_MODEL / AGENT_{NAME}_FALLBACK |
orchestrator: gemini; others: ollama; fallback: none |
Per-agent provider and optional fallback (gemini | ollama | openai_compatible). Fallback targets are not restricted; choose them according to the deployment's data-locality requirements. |
EDGE_LOCAL_ONLY |
false |
When true, roles other than the orchestrator resolve only to local providers; a cloud default or fallback for an edge role is treated as unavailable. |
DEFAULT_TEMPERATURE |
0.7 |
Generation temperature. |
DECISION_TEMPERATURE |
0.1 |
Low temperature for autonomous agent decisions (rule adherence on small models). |
MAX_STEP_RETRIES |
2 |
In-place retries per step for transient errors. |
MAX_REPLANS |
1 |
Bounded re-planning rounds for parameter errors. |
Per-agent serving endpoints can also be set in code via
AgentCapability(endpoint="http://my-slm-host:8000/v1", ...), which overrides
the global OpenAI-compatible base URL for that agent.
MACE/
├── mace/
│ ├── __init__.py # Public API exports
│ ├── state.py # AgentState (shared workflow state)
│ ├── settings.py # Configuration (pydantic-settings)
│ ├── registry.py # AgentRegistry, AgentCapability (+ endpoint)
│ ├── llm_config.py # LLM factories + per-agent fallback chain
│ ├── llm_utils.py # Robust JSON parsing, content extraction
│ ├── sanitization_bridge.py # Layer 1: whitelist reconstruction
│ ├── safety_agent.py # Layer 2: local semantic compression
│ ├── orchestrator.py # Cloud planner + checkpoint evaluation
│ ├── step_executor.py # Retries, re-planning, re-delegation
│ ├── base_agent.py # BaseEdgeAgent (+ delegate / ask_user)
│ ├── auth.py # Identity token threading hooks
│ ├── security.py # Log masking utilities
│ └── service.py # MACEService (LangGraph wiring)
├── examples/quickstart.py # Runnable demo (offline-capable)
├── docker-compose.yml # Ollama + demo stack
├── Dockerfile
├── pyproject.toml
├── .env.example
├── LICENSE # MIT
└── README.md
This repository is the open-source extraction of the MACE (Multi-Agent Cloud-Edge) framework described in an academic study of privacy-preserving hybrid agent architectures. The production system from which it was extracted applies MACE to an organisational data-management deployment with domain-specific agents (database/API operation, knowledge retrieval, report generation, calculation, analysis) on top of exactly the mechanisms published here.
Citation: [Paper under review; citation to be added]
- Research prototype. This code base prioritises architectural clarity over production hardening. It has not been security-audited; use it as a reference implementation.
- Layer 2 is best-effort. The Safety Agent uses a local LLM over untrusted text and can miss or over-redact content; Layer 1 (template reconstruction) is the hard boundary and only covers structured results. Free-text channels ultimately rely on Layer 2 plus pattern rules.
- Templates are deployment-specific. The fourteen default fuzzification templates are domain-neutral equivalents of the fourteen domain-specific templates used in the production deployment; real deployments must define their own action keys and sensitive-field lists.
- The Orchestrator sees metadata. Record counts, status flags, error categories, task descriptions, and the user's own messages do reach the cloud. MACE bounds what leaks (structure, not content); it is not a formal information-flow guarantee.
- Sequential execution. The step executor runs plan steps sequentially; there is no parallel step scheduling.
- In-memory state. Conversation and plan state live in process memory; persistence, streaming progress events, and permission middleware from the production deployment are not included in this extraction.
- LLM-dependent planning quality. Plan and checkpoint quality depend on the cloud model; malformed plans are retried and fail closed to a general-chat answer, but no formal plan validation is performed.
Issues and pull requests are welcome.
Cloud-bound semantic sanitization requires an available local model. Short text uses the same model path as long text. Successful output may be truncated; unavailability, call failure, or empty output raises SanitizationError. Checkpoint, conversation-history, and delegation callers propagate this failure before issuing the dependent cloud request. Agent processors raising this error are not retried or converted into ordinary task failures. Applications should catch SanitizationError at the request boundary, report a local failure, and stop that request. Do not resume with the original text. Independent rule filtering and local user-response template filling remain separate utilities, not fallbacks for cloud-bound semantic sanitization.
Regression tests use simulated model responses and failures, without cloud calls; they verify control flow, not real-model semantic sanitization quality.
Safety Agent automatic resolution uses only its configured primary local provider; known cloud providers and provider fallback are disabled for this role. Custom OpenAI-compatible endpoints and explicitly injected model objects must point to local infrastructure; the framework cannot infer the trust boundary of arbitrary URLs.