Harden edge0 serve: loopback bind warning and chat request limits - #4
Open
inputdrive wants to merge 1 commit into
Open
inputdrive wants to merge 1 commit into
inputdrive wants to merge 1 commit into
Conversation
Default bind remains 127.0.0.1; warn on non-loopback hosts that would expose the unauthenticated OpenAI-compatible API. Cap request bodies at 1 MiB (413), clamp max_tokens to 2048, and reject oversized message lists/prompts. Document residual Jinja chat_template and trust_remote_code risks in docs/HARDENING.md. Co-authored-by: inputdrive <inputdrive@gmail.com>
There was a problem hiding this comment.
🟡 Changes recommended
There are a couple of verified user-visible/robustness issues in the new hardening code (duplicate bind warning emission and a potential TypeError crash in multipart message handling) that should be fixed before approval.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
This PR hardens the unauthenticated edge0 serve OpenAI-compatible HTTP endpoints for safer local development by warning on non-loopback binds and enforcing request-shape/cost limits on the chat API.
Changes:
- Add bind-safety helpers and emit a warning when serving on non-loopback interfaces.
- Add chat request validation/limits (body size cap, max token clamp, message/prompt bounds) and ensure transports return consistent JSON errors.
- Add/extend tests and document the security posture and residual checkpoint-trust risks.
File summaries
| File | Description |
|---|---|
| tests/test_server.py | Adds unit/integration coverage for request limits, stdlib/Flask behaviors, and bind helpers. |
| tests/test_cli.py | Verifies CLI serve defaults and explicit host/port parsing. |
| src/edge0/server/limits.py | Introduces shared constants/helpers for request caps and loopback detection. |
| src/edge0/server/chat.py | Enforces message/prompt limits and clamps max_tokens during chat request parsing. |
| src/edge0/server/app.py | Applies request caps in both Flask and stdlib transports; normalizes JSON error responses and handler result serialization. |
| src/edge0/server/init.py | Exposes ChatRequestError in the server public API surface. |
| src/edge0/cli.py | Adds build_parser() for testability and prints bind warnings when serving. |
| README.md | Links new hardening documentation. |
| pyproject.toml | Documents rationale for the yanked mlx-lm==0.31.0 pin. |
| docs/models/edge0-8b.md | Notes serve hardening and links to HARDENING.md. |
| docs/models/edge0-35b.md | Notes serve hardening and links to HARDENING.md. |
| docs/HARDENING.md | Adds a dedicated hardening/security guidance document. |
Review details
- Files reviewed: 12/12 changed files
- Comments generated: 3
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+85
to
+87
| if isinstance(content, list): | ||
| return "".join( | ||
| p.get("text", "") for p in content if isinstance(p, dict)) |
Comment on lines
+195
to
+198
| warning = insecure_bind_warning(args.host, args.port) | ||
| if warning: | ||
| print(warning, file=sys.stderr) | ||
|
|
| # so localhost clients that send OpenAI-style large caps still work). | ||
| MAX_MAX_TOKENS = 2048 | ||
|
|
||
| # Chat-compleitions prompt shape. 400 when exceeded. |
Besho777
approved these changes
Sep 10, 2026
Besho777
approved these changes
Sep 10, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Harden
edge0 servefor local use without breaking the localhost OpenAI-compatible API.The HTTP surface (
/healthz,/v1/models,/v1/chat/completions) is unauthenticated. This PR makes the default bind harder to misuse and bounds request cost.Changes
--hoststays127.0.0.1. Explicit non-loopback binds (0.0.0.0, LAN IPs,::, hostnames) print a stderr warning that/v1/*has no auth.413(stdlibContent-Lengthcheck; FlaskMAX_CONTENT_LENGTH)max_tokensclamped to ≤ 2048 (localhost clients that send a large OpenAI-style cap still get200)messagesmust be an array of objects (avoids iterating a huge string)docs/HARDENING.md— localhost default, do not expose/v1/*without a proxy/auth, checkpoint trust, yankedmlx-lm==0.31.0./v1/completionsand validation errors return JSON400instead of crashing on a 2-tuple handler result.mlx-lm==0.31.0is not bumped. PyPI yanked it for batched KV-cache cross contamination; newer mlx-lm has historically broken edge0 decode (tolist()on lazy arrays). edge0 serializes generations (one request at a time), so the batch-cache bug is off the serving path. Seepyproject.tomland HARDENING.md.Residual risks (not fully mitigated here)
These are trusted-checkpoint issues, documented rather than sandboxed:
chat_template— Ling /edge0-8brenderschat_template.jinjawith Jinja2 (autoescape=False, notSandboxedEnvironment). A malicious template can execute Python duringencode_chat. Qwen paths use tokenizerapply_chat_template(also template-driven).trust_remote_code/auto_map— tokenizer load isAutoTokenizer.from_pretrained(..., local_files_only=True, trust_remote_code=True). A checkpointauto_mapcan run tokenizer code from that directory. Safetensors weights are mmap'd (not pickled), but tokenizer/config side files are not isolated.Only load checkpoints from a trusted source. This PR does not add API keys, TLS, rate limits, or a template sandbox.
Tests
Content-Length, 400 on oversized prompts/message lists (includingstream: truebefore SSE headers), loopback bind helpers, and--hostdefault.pytest tests/test_server.py tests/test_cli.py tests/test_repo_hygiene.py— 41 passed (Linux, no MLX; Flask installed).mlxwheels). macOS CI should still cover those; this PR does not change kernels or generation.Security rationale
A default or accidental
0.0.0.0bind would publish an unauthenticated completions API on the LAN. Unboundedmax_tokens, message lists, and body size let a local (or LAN) client pin GPU/RAM. Caps keep the local-dev API usable while making the dangerous bind noisy and bounding DoS on the chat path.