Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MACE: Multi-Agent Cloud-Edge Hybrid Framework

A privacy-preserving multi-agent framework, built on LangGraph, that separates cloud-based reasoning from edge-based execution through a two-layer Sanitization Bridge.

MACE lets an organisation use a commercial cloud LLM (e.g., Gemini) for what it is best at (intent analysis, multi-step planning, checkpoint evaluation, response-template authoring) while all tools, data operations, and response generation run on local infrastructure with small language models (Ollama or any OpenAI-compatible endpoint). The two sides communicate exclusively through operation commands (cloud to edge) and fuzzy status descriptions (edge to cloud). Raw data never crosses the bridge.

Features

Feature Description
Two-layer Sanitization Bridge Layer 1: rule-based whitelist reconstruction of structured results via fuzzification templates (no LLM, injection-immune, fail-closed). Layer 2: every nonempty text uses local-model semantic sanitization, regardless of length; an unavailable model, failed call, or empty output raises SanitizationError and aborts the workflow. No rule-only fallback is used.
Cloud-edge hybrid LLM Orchestrator runs on a cloud LLM; edge agents and the Safety Agent run on local SLMs.
Dynamic agent registry Agents register capability descriptions (and optional per-agent serving endpoints) at runtime; the Orchestrator discovers them from the registry, never from hard-coded lists.
Plan-embedded checkpoints The Orchestrator inserts checkpoint steps into its plans; at each checkpoint it evaluates sanitized progress summaries and decides continue / stop / modify (with replacement steps).
Chained peer re-delegation An edge agent that cannot finish a task delegates it to a peer; the executor guards against self-delegation and routes duplicate delegations through an adjudication checkpoint. Peers may delegate further.
Fault tolerance Transient errors (HTTP 408/429/5xx) retry in place; parameter errors route back to the Orchestrator for bounded re-planning with the failure reason attached.
Per-agent LLM config with fallback Each agent has a default provider and an optional fallback (gemini / ollama / openai_compatible), configurable per agent via environment variables. The Safety Agent never uses provider fallback and rejects known cloud providers; other roles do not restrict fallback targets; no fallback is applied unless one is configured. EDGE_LOCAL_ONLY restricts every role except the orchestrator to local providers.
Identity token threading The end user's identity token is threaded through state to edge agents, which attach it to internal API calls (Authorization: Bearer ...). The framework holds no credential of its own and the token never enters a cloud prompt. Provider-agnostic (no identity-provider dependency).
Secure logging Optional log masking of sensitive fields and patterns (tokens, emails, card numbers).

Architecture

+------------------------------------------------------------------+
|                         User / Frontend                          |
|          (message + end-user identity token)                     |
+---------------------------------+--------------------------------+
                                  |
                                  v
+------------------------------------------------------------------+
|              Cloud Strategy Layer  (Orchestrator)                 |
|                                                                  |
|  - Intent analysis, multi-step planning     Cloud LLM (Gemini /  |
|  - Checkpoint evaluation (continue/         OpenAI-compatible)   |
|    stop/modify)                                                  |
|  - Response-template authoring              NEVER sees raw data  |
|                                             NEVER sees the token |
+---------------------------------+--------------------------------+
        commands, templates |            ^  fuzzy status descriptions
                            v            |
              +-----------------------------------+
              |        Sanitization Bridge        |
              |  L1: whitelist reconstruction     |
              |      (templates, no LLM,          |
              |      fail-closed)                 |
              |  L2: Safety Agent (local SLM      |
              |      semantic compression +       |
              |      pattern filtering)           |
              +-----------------------------------+
                            |            ^
                            v            |  raw results
+------------------------------------------------------------------+
|                Edge Execution Layer  (trusted side)               |
|                                                                  |
|  Step Executor: retries, re-planning, dynamic task insertion,    |
|                 checkpoint routing, context injection            |
|                                                                  |
|   +-----------+   +-----------+   +-----------+                  |
|   | Agent A   |   | Agent B   |   | Agent C   |   ...            |
|   | (queries) |   | (reports) |   | (calc)    |                  |
|   +-----------+   +-----------+   +-----------+                  |
|                                                                  |
|  - Full access to private data      - Local SLM (Ollama /        |
|  - Internal API calls made with       OpenAI-compatible)         |
|    the user's own token             - Final responses generated  |
|                                       on the edge                |
+------------------------------------------------------------------+

Core principle: logic-data separation

Cloud (Orchestrator) Edge (agents)
Task planning Access private data
Intent analysis Execute tools and API calls
Checkpoint evaluation Run calculations
Template authoring Generate detailed responses
No data access, no tool execution, no user token Holds the full reality

Quick start

git clone https://github.com/watsonshih/MACE.git
cd MACE
pip install -e .            # core only; dry-runs offline
python examples/quickstart.py

The quickstart runs a full pipeline of plan, edge execution, sanitization (raw vs. sanitized shown side by side), checkpoint decision, and final response, against an in-memory toy inventory. No cloud key is required: without GEMINI_API_KEY (or with MACE_OFFLINE=1) a deterministic offline planner stands in for the cloud LLM.

To use the real cloud Orchestrator and a local edge model:

pip install -e ".[gemini,ollama]"
cp .env.example .env         # fill in GEMINI_API_KEY
ollama pull qwen3.5-9b
python examples/quickstart.py

Or with Docker:

docker compose up --build

Minimal usage

import asyncio
from mace import (
    MACEService, BaseEdgeAgent, AgentCapability, AgentModelConfig,
    AgentState, SanitizationBridge, bearer_headers,
)

class MyDataAgent(BaseEdgeAgent):
    async def process(self, state: AgentState) -> AgentState:
        task = self.get_task_description(state)
        headers = bearer_headers(state)          # user's token, pass-through
        data = await my_internal_api.search(task, headers=headers)
        self.set_result(state, success=True, data=data,
                        message=f"Found {len(data)} records",
                        action_type="list_records")
        return state

bridge = SanitizationBridge(sensitive_fields={"salary", "revenue"})
service = MACEService(bridge=bridge)
service.register_agent(
    MyDataAgent(name="data_agent"),
    AgentCapability(
        name="data_agent",
        display_name="Data Agent",
        description="Queries the local database for records.",
        capabilities=["search_records", "get_record_by_id"],
        model_config=AgentModelConfig(default_model="ollama"),
    ),
)

async def main():
    print(await service.chat("Find all records from last month",
                             user_token=end_user_token))

asyncio.run(main())

Sanitization example

Raw edge result (stays on the edge):

{
  "success": true,
  "status": "success",
  "data": [
    {"item": "Widget B", "stock": 8, "unit_cost": 47.0,
     "supplier_email": "orders@supplier-b.example"},
    {"item": "Gadget C", "stock": 3, "unit_cost": 230.0,
     "supplier_email": "contact@supplier-c.example"}
  ],
  "message": "Found 2 items below their reorder point",
  "action_type": "stock_query"
}

What the cloud Orchestrator receives instead:

{
  "status": "success",
  "fuzzy_description": "Stock query returned 2 items",
  "data_available": true,
  "record_count": 2,
  "has_error": false
}

Layer 1 does not filter the raw payload; it reconstructs a brand-new object from a fixed schema and a whitelisted template. The only values derived from the raw result are a status flag, an integer count, an optional numeric rate, and (on failure) an error category from a fixed enumeration. Prompt injection payloads embedded in tool output therefore cannot survive the bridge, and an unrecognised action_type fails closed to a status-only summary ("Operation completed").

The default template set contains fourteen action keys (list_records, get_record, record_created, list_related_records, list_entries, entry_created, search_result, knowledge_query, calculation_completed, report_generated, completeness_check, external_fetch, error, default), mirroring the fourteen registered by the production deployment one-for-one. Production keys were domain-specific (several were per-resource instances of the same list/create archetypes); the defaults here are their domain-neutral equivalents, and applications are expected to register their own per-action templates via fuzzy_templates.

Configuration reference

All settings come from environment variables or a .env file (see .env.example).

Variable Default Description
GEMINI_API_KEY (none) Cloud LLM key (planning). Optional: without it, the orchestrator uses its configured fallback, if any.
GEMINI_MODEL gemini-3-flash-preview Cloud model name.
GEMINI_TIMEOUT / GEMINI_MAX_RETRIES 20 / 5 Per-call timeout (s) and SDK retries.
OLLAMA_BASE_URL http://localhost:11434 Edge LLM server.
OLLAMA_MODEL qwen3.5-9b Edge model name.
OLLAMA_NUM_CTX 128000 Edge context window (max tokens).
OPENAI_COMPATIBLE_BASE_URL / _API_KEY / _MODEL (none) Any OpenAI-compatible endpoint (vLLM, llama.cpp, MLX serving, ...).
AGENT_{NAME}_MODEL / AGENT_{NAME}_FALLBACK orchestrator: gemini; others: ollama; fallback: none Per-agent provider and optional fallback (gemini | ollama | openai_compatible). Fallback targets are not restricted; choose them according to the deployment's data-locality requirements.
EDGE_LOCAL_ONLY false When true, roles other than the orchestrator resolve only to local providers; a cloud default or fallback for an edge role is treated as unavailable.
DEFAULT_TEMPERATURE 0.7 Generation temperature.
DECISION_TEMPERATURE 0.1 Low temperature for autonomous agent decisions (rule adherence on small models).
MAX_STEP_RETRIES 2 In-place retries per step for transient errors.
MAX_REPLANS 1 Bounded re-planning rounds for parameter errors.

Per-agent serving endpoints can also be set in code via AgentCapability(endpoint="http://my-slm-host:8000/v1", ...), which overrides the global OpenAI-compatible base URL for that agent.

Project structure

MACE/
├── mace/
│   ├── __init__.py             # Public API exports
│   ├── state.py                # AgentState (shared workflow state)
│   ├── settings.py             # Configuration (pydantic-settings)
│   ├── registry.py             # AgentRegistry, AgentCapability (+ endpoint)
│   ├── llm_config.py           # LLM factories + per-agent fallback chain
│   ├── llm_utils.py            # Robust JSON parsing, content extraction
│   ├── sanitization_bridge.py  # Layer 1: whitelist reconstruction
│   ├── safety_agent.py         # Layer 2: local semantic compression
│   ├── orchestrator.py         # Cloud planner + checkpoint evaluation
│   ├── step_executor.py        # Retries, re-planning, re-delegation
│   ├── base_agent.py           # BaseEdgeAgent (+ delegate / ask_user)
│   ├── auth.py                 # Identity token threading hooks
│   ├── security.py             # Log masking utilities
│   └── service.py              # MACEService (LangGraph wiring)
├── examples/quickstart.py      # Runnable demo (offline-capable)
├── docker-compose.yml          # Ollama + demo stack
├── Dockerfile
├── pyproject.toml
├── .env.example
├── LICENSE                     # MIT
└── README.md

Relation to the paper

This repository is the open-source extraction of the MACE (Multi-Agent Cloud-Edge) framework described in an academic study of privacy-preserving hybrid agent architectures. The production system from which it was extracted applies MACE to an organisational data-management deployment with domain-specific agents (database/API operation, knowledge retrieval, report generation, calculation, analysis) on top of exactly the mechanisms published here.

Citation: [Paper under review; citation to be added]

Limitations

  • Research prototype. This code base prioritises architectural clarity over production hardening. It has not been security-audited; use it as a reference implementation.
  • Layer 2 is best-effort. The Safety Agent uses a local LLM over untrusted text and can miss or over-redact content; Layer 1 (template reconstruction) is the hard boundary and only covers structured results. Free-text channels ultimately rely on Layer 2 plus pattern rules.
  • Templates are deployment-specific. The fourteen default fuzzification templates are domain-neutral equivalents of the fourteen domain-specific templates used in the production deployment; real deployments must define their own action keys and sensitive-field lists.
  • The Orchestrator sees metadata. Record counts, status flags, error categories, task descriptions, and the user's own messages do reach the cloud. MACE bounds what leaks (structure, not content); it is not a formal information-flow guarantee.
  • Sequential execution. The step executor runs plan steps sequentially; there is no parallel step scheduling.
  • In-memory state. Conversation and plan state live in process memory; persistence, streaming progress events, and permission middleware from the production deployment are not included in this extraction.
  • LLM-dependent planning quality. Plan and checkpoint quality depend on the cloud model; malformed plans are retried and fail closed to a general-chat answer, but no formal plan validation is performed.

Contributing

Issues and pull requests are welcome.

License

MIT

Layer-2 failure policy

Cloud-bound semantic sanitization requires an available local model. Short text uses the same model path as long text. Successful output may be truncated; unavailability, call failure, or empty output raises SanitizationError. Checkpoint, conversation-history, and delegation callers propagate this failure before issuing the dependent cloud request. Agent processors raising this error are not retried or converted into ordinary task failures. Applications should catch SanitizationError at the request boundary, report a local failure, and stop that request. Do not resume with the original text. Independent rule filtering and local user-response template filling remain separate utilities, not fallbacks for cloud-bound semantic sanitization.

Regression tests use simulated model responses and failures, without cloud calls; they verify control flow, not real-model semantic sanitization quality.

Safety Agent automatic resolution uses only its configured primary local provider; known cloud providers and provider fallback are disabled for this role. Custom OpenAI-compatible endpoints and explicitly injected model objects must point to local infrastructure; the framework cannot infer the trust boundary of arbitrary URLs.

About

Multi-Agent Cloud-Edge: a privacy-preserving hybrid framework with a two-layer Sanitization Bridge

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages