Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 28 additions & 3 deletions manual/english/Searching/Conversational_search.md
Original file line number Diff line number Diff line change
Expand Up @@ -108,8 +108,8 @@ You can also set provider options and retrieval limits:
```sql
CREATE CHAT MODEL support_assistant (
model='openai:gpt-4o-mini',
api_key='your-provider-api-key',
base_url='http://host.docker.internal:8787/v1',
api_key='your-azure-entra-token',
base_url='https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/responses',
timeout=60,
retrieval_limit=5,
max_document_length=3000
Expand All @@ -125,7 +125,7 @@ CREATE CHAT MODEL support_assistant (
```bash
curl -s -X POST 'http://localhost:9308/sql?mode=raw' \
-H 'Content-Type: text/plain' \
-d "CREATE CHAT MODEL support_assistant (model='openai:gpt-4o-mini', api_key='your-provider-api-key', base_url='http://host.docker.internal:8787/v1', timeout=60, retrieval_limit=5, max_document_length=3000)"
-d "CREATE CHAT MODEL support_assistant (model='openai:gpt-4o-mini', api_key='your-azure-entra-token', base_url='https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/responses', timeout=60, retrieval_limit=5, max_document_length=3000)"
```

<!-- end -->
Expand Down Expand Up @@ -161,6 +161,31 @@ environment:

If `api_key` is not set in `CREATE CHAT MODEL`, the `llm` extension can use the matching provider environment variable. Set `api_key` in the chat model only when you need this model to use a different key.

## Local OpenAI-compatible models

LM Studio can run local models with Conversational Search. For a local model identifier that is not in the `openai:*` catalog, use the `openrouter:*` transport with LM Studio's Chat Completions endpoint. When Buddy runs in Docker, use `host.docker.internal` to reach the server on the host.

<!-- example conversational_search_create_local_model -->

<!-- intro -->
##### SQL:

<!-- request SQL -->

```sql
CREATE CHAT MODEL local_assistant (
model='openrouter:google_gemma-4-e4b-it',
api_key='lm-studio',
base_url='http://host.docker.internal:1234/v1/chat/completions',
timeout=60,
retrieval_limit=5
);
```

<!-- end -->

`api_key` is a non-secret placeholder accepted by LM Studio; do not use a real provider key for a local server. The `llm` extension validates the model portion of `openai:*` against its supported OpenAI model names, so `openai:google_gemma-4-e4b-it` is rejected even when LM Studio has that model loaded. The `openrouter:*` transport accepts the local identifier and can target an OpenAI-compatible Chat Completions server through `base_url`. Conversational Search requires the local model to reliably return an OpenAI function call for Buddy's routing schema. Test it with `CALL CHAT`; basic text completion or a simple tool-call test alone is not sufficient.

## CALL CHAT syntax

```sql
Expand Down
49 changes: 49 additions & 0 deletions test/clt-tests/buddy-plugins/test-conversational-basic.rec
Original file line number Diff line number Diff line change
Expand Up @@ -191,6 +191,55 @@ mysql -h0 -P9306 -e "CALL CHAT('Can you explain more?', 'docs', 'test_assistant'
#!/.*response: .+/!#
#!/.*sources: \[\]/!#
––– comment –––
Create a dedicated chat model for the endpoint probe. It is intentionally
allowed to receive HTTP 502 responses and is never used by normal chat cases.
––– input –––
rm -f /tmp/chat-model-requests.log /tmp/mock-chat-model-server.php
cat > /tmp/mock-chat-model-server.php <<'PHP'
<?php
$logFile = getenv('CHAT_MODEL_REQUEST_LOG') ?: '';
if ($logFile !== '') {
file_put_contents($logFile, ($_SERVER['REQUEST_URI'] ?? '') . PHP_EOL, FILE_APPEND | LOCK_EX);
}
http_response_code(502);
header('Content-Type: application/json');
echo '{"error":{"message":"CLT endpoint probe"}}';
PHP
CHAT_MODEL_REQUEST_LOG=/tmp/chat-model-requests.log php -S 127.0.0.1:18087 /tmp/mock-chat-model-server.php > /tmp/mock-chat-model-server.log 2>&1 &
MOCK_CHAT_PID=$!
sleep 1
mysql -h0 -P9306 << 'EOF'
CREATE CHAT MODEL endpoint_probe_assistant (
model='openai:gpt-4o-mini',
api_key='chat-model-key',
base_url='http://127.0.0.1:18087/configured/v1/responses',
timeout=60
);
EOF
––– output –––
+--------------------------+
| name |
+--------------------------+
| endpoint_probe_assistant |
+--------------------------+
––– comment –––
Test CALL CHAT uses the dedicated model-specific base_url rather than the environment default
––– input –––
mysql -h0 -P9306 -e "CALL CHAT('What is vector search?', 'docs', 'endpoint_probe_assistant')" > /dev/null 2>&1 || true
sort -u /tmp/chat-model-requests.log
––– output –––
/configured/v1/responses
––– comment –––
Stop the local endpoint probe
––– input –––
kill "$MOCK_CHAT_PID" 2>/dev/null || true
––– output –––
––– comment –––
Remove the dedicated endpoint-probe model
––– input –––
mysql -h0 -P9306 -e "DROP CHAT MODEL endpoint_probe_assistant;"
––– output –––
––– comment –––
Test DROP CHAT MODEL - cleanup
––– input –––
mysql -h0 -P9306 -e "DROP CHAT MODEL test_assistant;"
Expand Down