Skip to content

feat(server): OpenAI-compatible tool/function calling - #119

Open
agourakis82 wants to merge 1 commit into
Edge0-AI:mainfrom
agourakis82:feat/openai-tool-calling
Open

agourakis82 wants to merge 1 commit into
Edge0-AI:mainfrom
agourakis82:feat/openai-tool-calling

Conversation

@agourakis82

Copy link
Copy Markdown
Contributor

Fixes #115.

The server ignored tools/tool_choice in chat-completion requests and never returned tool_calls, even though both checkpoints' chat templates support tool calling and the model can already produce the expected XML on request.

Request side

  • ChatRequest gains tools/tool_choice; ChatMessage gains tool_calls/tool_call_id/name for multi-turn tool round trips (issue's ask Thank you for creating this model. How do I run it on my phone? #5).
  • ChatSession.prompt_ids() forwards tool definitions to both template paths: Ling8BEngine.encode_chat's tools= kwarg was already wired to the template but hardcoded to None — this was the "partial plumbing" that made the fix smaller than it looked. The qwen path now passes tools= into tok.apply_chat_template().
  • tool_choice: "none" withholds tool definitions from the prompt. Forcing/targeting a specific function (tool_choice: "required" or {"type":"function",...}) would need constrained decoding — out of scope here, both currently behave like "auto".
  • Found and fixed a latent bug while wiring this up: an assistant message with content: null (the OpenAI convention for a tool-call-only turn) was coerced to the literal string "None" — m.get("content", "")'s default doesn't apply when the key is present with value None.

Response side

  • New src/edge0/server/tool_calls.py parses each family's XML dialect back into OpenAI's message.tool_calls schema. The two checkpoints use two different dialects — confirmed against the real chat_template.jinja each ships, not assumed from the issue's example:
    • edge0-8b (Ling): <tool_call>{name}<arg_key>k</arg_key><arg_value>v</arg_value>...</tool_call>
    • edge0-35b (Qwen3.5): <tool_call>\n<function={name}>\n<parameter={key}>\nvalue\n</parameter>\n...</function></tool_call>
      Each engine gets a parse_tool_calls() method mirroring the existing per-engine encode_chat pattern, so app.py stays dialect-agnostic.
  • _chat_once: strips the XML block from content, sets content: null and finish_reason: "tool_calls" when a call is found — matching the issue's expected-result example exactly. Unchanged when the request has no tools or the engine has no parser.
  • _chat_stream: tool-enabled requests buffer the full response and parse once, emitting tool_calls only in the final chunk. Fully incremental per-argument SSE deltas would need token-level structured buffering this codebase doesn't have anywhere yet (reasoning_content isn't split incrementally in streaming either — same existing asymmetry). Requests without tools keep the current immediate per-token streaming, unchanged — zero behavior change for the common case.

Verification

This machine happens to have both real checkpoints locally, so I didn't stop at synthetic fixtures:

  • encode_chat(messages, tools=[calculator]) against the real edge0-8b checkpoint renders the actual tool-prompt end to end (236 tokens, correct # Tools / <tool_call> format section).
  • A live generation at temperature=0 against that same checkpoint produced a genuine <tool_call>calculator\n<arg_key>expression</arg_key>\n<arg_value>2 + 2</arg_value>\n</tool_call> block, which parse_ling_tool_calls correctly extracted into {"function": {"name": "calculator", "arguments": "{\"expression\": \"2 + 2\"}"}} — the real model, the real template, the real parser, not a mock.
  • The Qwen3.5 dialect's test fixtures are the real 35B template's own in-prompt format example (example_function_name/multi-line parameter value), copied verbatim from a jinja2 render of the real chat_template.jinja.

Test plan

  • tests/test_tool_calls.py (new): both XML dialects, multiple calls in one response, prose-before-call preservation, malformed-block-preserved-as-content, argument type coercion (numbers/bools/strings).
  • tests/test_server.py: tools/tool_choice parsing, content: null round trip, _chat_once/_chat_stream tool-call extraction (including a regression guard that no-tools requests keep the existing immediate per-token SSE streaming, unchanged).
  • pytest -m 'not slow': 94 passed, 1 skipped (up from the 79-test baseline; no regressions).

Not done (flagged, not silently skipped)

  • Forcing a specific function via tool_choice.
  • Fully incremental SSE tool-call argument streaming (buffer-then-emit-once is the current design, consistent with how reasoning_content is already handled in streaming).

The server ignored `tools`/`tool_choice` in chat-completion requests
and never returned `tool_calls`, even though both checkpoints' chat
templates support tool calling and the model can already produce the
expected XML on request.

## Request side
- `ChatRequest` gains `tools`/`tool_choice`; `ChatMessage` gains
  `tool_calls`/`tool_call_id`/`name` for multi-turn tool round trips.
- `ChatSession.prompt_ids()` forwards tool definitions to both template
  paths: `Ling8BEngine.encode_chat`'s `tools=` kwarg was already wired
  to the template but hardcoded to `None`; the qwen path now passes
  `tools=` into `tok.apply_chat_template()`. `tool_choice: "none"`
  withholds tool definitions (no forced/targeted call -- that needs
  constrained decoding, out of scope here).
- Fixed a latent bug found while wiring this up: `content: null` on an
  assistant tool-call message was coerced to the literal string
  `"None"` (`m.get("content", "")` doesn't apply the default when the
  key is present with value `None`).

## Response side
- New `edge0/server/tool_calls.py` parses each family's XML dialect
  back into OpenAI's `message.tool_calls` schema. The two checkpoints
  use two different dialects -- confirmed against the real
  `chat_template.jinja` each ships (not assumed):
  - edge0-8b (Ling): `<tool_call>{name}<arg_key>k</arg_key>
    <arg_value>v</arg_value>...</tool_call>`
  - edge0-35b (Qwen3.5): `<tool_call>\n<function={name}>\n
    <parameter={key}>\nvalue\n</parameter>\n...</function></tool_call>`
  Each engine gets a `parse_tool_calls()` method (mirroring the
  existing per-engine `encode_chat` pattern) so `app.py` stays
  dialect-agnostic.
- `_chat_once`: strips the XML block from `content`, sets `content` to
  `None` and `finish_reason` to `"tool_calls"` when a call is found;
  unchanged when the request has no `tools` or the engine has no
  parser.
- `_chat_stream`: tool-enabled requests now buffer the full response
  and parse once, emitting `tool_calls` only in the final chunk --
  incrementally streaming per-argument JSON deltas would need
  token-level structured buffering this codebase doesn't have
  anywhere yet (reasoning_content isn't split in streaming either).
  Requests without `tools` keep the existing immediate per-token
  streaming, unchanged.

## Verification
Validated end to end against the real edge0-8b checkpoint, not just
fixtures: `encode_chat(..., tools=[...])` renders the real tool-prompt,
and a live generation at temperature=0 produced a genuine
`<tool_call>calculator<arg_key>expression</arg_key>...` block that
`parse_ling_tool_calls` correctly extracted into
`{"expression": "2 + 2"}`.

Not done (flagged, not silently skipped): forcing a specific function
via `tool_choice`, and fully incremental SSE tool-call argument
streaming. Both tool_choice: "auto" and the full non-streaming/
buffered-streaming round trip work.

Fixes Edge0-AI#115.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

OpenAI-compatible server ignores tools and never returns tool_calls

1 participant