feat(server): OpenAI-compatible tool/function calling - #119
Open
agourakis82 wants to merge 1 commit into
Open
agourakis82 wants to merge 1 commit into
agourakis82 wants to merge 1 commit into
Conversation
The server ignored `tools`/`tool_choice` in chat-completion requests
and never returned `tool_calls`, even though both checkpoints' chat
templates support tool calling and the model can already produce the
expected XML on request.
## Request side
- `ChatRequest` gains `tools`/`tool_choice`; `ChatMessage` gains
`tool_calls`/`tool_call_id`/`name` for multi-turn tool round trips.
- `ChatSession.prompt_ids()` forwards tool definitions to both template
paths: `Ling8BEngine.encode_chat`'s `tools=` kwarg was already wired
to the template but hardcoded to `None`; the qwen path now passes
`tools=` into `tok.apply_chat_template()`. `tool_choice: "none"`
withholds tool definitions (no forced/targeted call -- that needs
constrained decoding, out of scope here).
- Fixed a latent bug found while wiring this up: `content: null` on an
assistant tool-call message was coerced to the literal string
`"None"` (`m.get("content", "")` doesn't apply the default when the
key is present with value `None`).
## Response side
- New `edge0/server/tool_calls.py` parses each family's XML dialect
back into OpenAI's `message.tool_calls` schema. The two checkpoints
use two different dialects -- confirmed against the real
`chat_template.jinja` each ships (not assumed):
- edge0-8b (Ling): `<tool_call>{name}<arg_key>k</arg_key>
<arg_value>v</arg_value>...</tool_call>`
- edge0-35b (Qwen3.5): `<tool_call>\n<function={name}>\n
<parameter={key}>\nvalue\n</parameter>\n...</function></tool_call>`
Each engine gets a `parse_tool_calls()` method (mirroring the
existing per-engine `encode_chat` pattern) so `app.py` stays
dialect-agnostic.
- `_chat_once`: strips the XML block from `content`, sets `content` to
`None` and `finish_reason` to `"tool_calls"` when a call is found;
unchanged when the request has no `tools` or the engine has no
parser.
- `_chat_stream`: tool-enabled requests now buffer the full response
and parse once, emitting `tool_calls` only in the final chunk --
incrementally streaming per-argument JSON deltas would need
token-level structured buffering this codebase doesn't have
anywhere yet (reasoning_content isn't split in streaming either).
Requests without `tools` keep the existing immediate per-token
streaming, unchanged.
## Verification
Validated end to end against the real edge0-8b checkpoint, not just
fixtures: `encode_chat(..., tools=[...])` renders the real tool-prompt,
and a live generation at temperature=0 produced a genuine
`<tool_call>calculator<arg_key>expression</arg_key>...` block that
`parse_ling_tool_calls` correctly extracted into
`{"expression": "2 + 2"}`.
Not done (flagged, not silently skipped): forcing a specific function
via `tool_choice`, and fully incremental SSE tool-call argument
streaming. Both tool_choice: "auto" and the full non-streaming/
buffered-streaming round trip work.
Fixes Edge0-AI#115.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #115.
The server ignored
tools/tool_choicein chat-completion requests and never returnedtool_calls, even though both checkpoints' chat templates support tool calling and the model can already produce the expected XML on request.Request side
ChatRequestgainstools/tool_choice;ChatMessagegainstool_calls/tool_call_id/namefor multi-turn tool round trips (issue's ask Thank you for creating this model. How do I run it on my phone? #5).ChatSession.prompt_ids()forwards tool definitions to both template paths:Ling8BEngine.encode_chat'stools=kwarg was already wired to the template but hardcoded toNone— this was the "partial plumbing" that made the fix smaller than it looked. The qwen path now passestools=intotok.apply_chat_template().tool_choice: "none"withholds tool definitions from the prompt. Forcing/targeting a specific function (tool_choice: "required"or{"type":"function",...}) would need constrained decoding — out of scope here, both currently behave like"auto".content: null(the OpenAI convention for a tool-call-only turn) was coerced to the literal string"None"—m.get("content", "")'s default doesn't apply when the key is present with valueNone.Response side
src/edge0/server/tool_calls.pyparses each family's XML dialect back into OpenAI'smessage.tool_callsschema. The two checkpoints use two different dialects — confirmed against the realchat_template.jinjaeach ships, not assumed from the issue's example:<tool_call>{name}<arg_key>k</arg_key><arg_value>v</arg_value>...</tool_call><tool_call>\n<function={name}>\n<parameter={key}>\nvalue\n</parameter>\n...</function></tool_call>Each engine gets a
parse_tool_calls()method mirroring the existing per-engineencode_chatpattern, soapp.pystays dialect-agnostic._chat_once: strips the XML block fromcontent, setscontent: nullandfinish_reason: "tool_calls"when a call is found — matching the issue's expected-result example exactly. Unchanged when the request has notoolsor the engine has no parser._chat_stream: tool-enabled requests buffer the full response and parse once, emittingtool_callsonly in the final chunk. Fully incremental per-argument SSE deltas would need token-level structured buffering this codebase doesn't have anywhere yet (reasoning_content isn't split incrementally in streaming either — same existing asymmetry). Requests withouttoolskeep the current immediate per-token streaming, unchanged — zero behavior change for the common case.Verification
This machine happens to have both real checkpoints locally, so I didn't stop at synthetic fixtures:
encode_chat(messages, tools=[calculator])against the real edge0-8b checkpoint renders the actual tool-prompt end to end (236 tokens, correct# Tools/<tool_call>format section).<tool_call>calculator\n<arg_key>expression</arg_key>\n<arg_value>2 + 2</arg_value>\n</tool_call>block, whichparse_ling_tool_callscorrectly extracted into{"function": {"name": "calculator", "arguments": "{\"expression\": \"2 + 2\"}"}}— the real model, the real template, the real parser, not a mock.example_function_name/multi-line parameter value), copied verbatim from a jinja2 render of the realchat_template.jinja.Test plan
tests/test_tool_calls.py(new): both XML dialects, multiple calls in one response, prose-before-call preservation, malformed-block-preserved-as-content, argument type coercion (numbers/bools/strings).tests/test_server.py:tools/tool_choiceparsing,content: nullround trip,_chat_once/_chat_streamtool-call extraction (including a regression guard that no-tools requests keep the existing immediate per-token SSE streaming, unchanged).pytest -m 'not slow': 94 passed, 1 skipped (up from the 79-test baseline; no regressions).Not done (flagged, not silently skipped)
tool_choice.