Skip to content

Harden edge0 serve: loopback bind warning and chat request limits - #4

Open
inputdrive wants to merge 1 commit into
Edge0-AI:mainfrom
inputdrive:cursor/security-hardening-serve-943c
Open

inputdrive wants to merge 1 commit into
Edge0-AI:mainfrom
inputdrive:cursor/security-hardening-serve-943c

Conversation

@inputdrive

Copy link
Copy Markdown

Summary

Harden edge0 serve for local use without breaking the localhost OpenAI-compatible API.

The HTTP surface (/healthz, /v1/models, /v1/chat/completions) is unauthenticated. This PR makes the default bind harder to misuse and bounds request cost.

Changes

  • Bind: CLI --host stays 127.0.0.1. Explicit non-loopback binds (0.0.0.0, LAN IPs, ::, hostnames) print a stderr warning that /v1/* has no auth.
  • Chat API limits:
    • Request body > 1 MiB → 413 (stdlib Content-Length check; Flask MAX_CONTENT_LENGTH)
    • max_tokens clamped to ≤ 2048 (localhost clients that send a large OpenAI-style cap still get 200)
    • 128 messages or > 100k prompt characters → 400

    • messages must be an array of objects (avoids iterating a huge string)
  • Docs: docs/HARDENING.md — localhost default, do not expose /v1/* without a proxy/auth, checkpoint trust, yanked mlx-lm==0.31.0.
  • Stdlib errors: /v1/completions and validation errors return JSON 400 instead of crashing on a 2-tuple handler result.

mlx-lm==0.31.0 is not bumped. PyPI yanked it for batched KV-cache cross contamination; newer mlx-lm has historically broken edge0 decode (tolist() on lazy arrays). edge0 serializes generations (one request at a time), so the batch-cache bug is off the serving path. See pyproject.toml and HARDENING.md.

Residual risks (not fully mitigated here)

These are trusted-checkpoint issues, documented rather than sandboxed:

  1. Jinja chat_template — Ling / edge0-8b renders chat_template.jinja with Jinja2 (autoescape=False, not SandboxedEnvironment). A malicious template can execute Python during encode_chat. Qwen paths use tokenizer apply_chat_template (also template-driven).
  2. trust_remote_code / auto_map — tokenizer load is AutoTokenizer.from_pretrained(..., local_files_only=True, trust_remote_code=True). A checkpoint auto_map can run tokenizer code from that directory. Safetensors weights are mmap'd (not pickled), but tokenizer/config side files are not isolated.

Only load checkpoints from a trusted source. This PR does not add API keys, TLS, rate limits, or a template sandbox.

Tests

  • Unit tests for clamp/reject behavior, 413 on oversize Content-Length, 400 on oversized prompts/message lists (including stream: true before SSE headers), loopback bind helpers, and --host default.
  • Ran pytest tests/test_server.py tests/test_cli.py tests/test_repo_hygiene.py — 41 passed (Linux, no MLX; Flask installed).
  • MLX math/registry/e2e tests were not run here (no Apple Silicon / mlx wheels). macOS CI should still cover those; this PR does not change kernels or generation.

Security rationale

A default or accidental 0.0.0.0 bind would publish an unauthenticated completions API on the LAN. Unbounded max_tokens, message lists, and body size let a local (or LAN) client pin GPU/RAM. Caps keep the local-dev API usable while making the dangerous bind noisy and bounding DoS on the chat path.

Default bind remains 127.0.0.1; warn on non-loopback hosts that would
expose the unauthenticated OpenAI-compatible API. Cap request bodies at
1 MiB (413), clamp max_tokens to 2048, and reject oversized message
lists/prompts. Document residual Jinja chat_template and
trust_remote_code risks in docs/HARDENING.md.

Co-authored-by: inputdrive <inputdrive@gmail.com>
Copilot AI lite review requested due to automatic review settings September 10, 2026 09:07

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

There are a couple of verified user-visible/robustness issues in the new hardening code (duplicate bind warning emission and a potential TypeError crash in multipart message handling) that should be fixed before approval.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

This PR hardens the unauthenticated edge0 serve OpenAI-compatible HTTP endpoints for safer local development by warning on non-loopback binds and enforcing request-shape/cost limits on the chat API.

Changes:

  • Add bind-safety helpers and emit a warning when serving on non-loopback interfaces.
  • Add chat request validation/limits (body size cap, max token clamp, message/prompt bounds) and ensure transports return consistent JSON errors.
  • Add/extend tests and document the security posture and residual checkpoint-trust risks.
File summaries
File Description
tests/test_server.py Adds unit/integration coverage for request limits, stdlib/Flask behaviors, and bind helpers.
tests/test_cli.py Verifies CLI serve defaults and explicit host/port parsing.
src/edge0/server/limits.py Introduces shared constants/helpers for request caps and loopback detection.
src/edge0/server/chat.py Enforces message/prompt limits and clamps max_tokens during chat request parsing.
src/edge0/server/app.py Applies request caps in both Flask and stdlib transports; normalizes JSON error responses and handler result serialization.
src/edge0/server/init.py Exposes ChatRequestError in the server public API surface.
src/edge0/cli.py Adds build_parser() for testability and prints bind warnings when serving.
README.md Links new hardening documentation.
pyproject.toml Documents rationale for the yanked mlx-lm==0.31.0 pin.
docs/models/edge0-8b.md Notes serve hardening and links to HARDENING.md.
docs/models/edge0-35b.md Notes serve hardening and links to HARDENING.md.
docs/HARDENING.md Adds a dedicated hardening/security guidance document.
Review details
  • Files reviewed: 12/12 changed files
  • Comments generated: 3
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +85 to +87
if isinstance(content, list):
return "".join(
p.get("text", "") for p in content if isinstance(p, dict))
Comment thread src/edge0/cli.py
Comment on lines +195 to +198
warning = insecure_bind_warning(args.host, args.port)
if warning:
print(warning, file=sys.stderr)

# so localhost clients that send OpenAI-style large caps still work).
MAX_MAX_TOKENS = 2048

# Chat-compleitions prompt shape. 400 when exceeded.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants