Skip to content

Commit 09a8c4e

Browse files
authored
fix(ai): only capture terminal, named Responses stop reasons (#919)
* fix(ai): only capture terminal, named Responses stop reasons The Python SDK recorded raw Responses API statuses as $ai_stop_reason across three surfaces, none of which matched what the trace view understands. The native wrapper returned any status, so a background run recorded queued as its stop reason, and a truncated run recorded the bare incomplete instead of max_output_tokens. The streaming accumulator only read status on response.completed, so streams that ended incomplete or failed recorded nothing. The LangChain callback read only generation_info.finish_reason, so Responses API and Anthropic runs, which report through response_metadata, recorded nothing at all. One shared _responses_stop_reason helper now does the mapping: non-terminal statuses yield nothing, an incomplete run is named by incomplete_details.reason, and the other terminal statuses stand for themselves. All three surfaces route through it, and the LangChain callback reads the same metadata sources as the JS SDK's callback, in the same priority order. Ports PostHog/posthog-js#4700 and PostHog/posthog-js#4736 to Python. Generated-By: PostHog Desktop Task-Id: 5cb23cd6-224c-4d86-bc24-97730d4c2f98 * chore(ai): tighten the stop reason helpers and their tests Same behavior, less of it. The metadata sources in the callback normalize to dicts once instead of each carrying its own isinstance guard, and the Responses helper folds its None guard and nested incomplete branch into the checks that already covered them. The tests drop a chat-completions passthrough case that test_openai.py already locks, hoist the imports the new cases share, and express the streaming behavior as a table rather than one test driving three accumulators, which also covers the failed status the old test missed. Generated-By: PostHog Desktop Task-Id: 5cb23cd6-224c-4d86-bc24-97730d4c2f98 * chore(ai): drop the metadata source Python langchain never populates The stop reason scan read response_metadata nested inside generation_info, mirroring the JS callback. Python langchain never nests it: generation_info carries finish_reason, logprobs, headers and gateway metadata, while response_metadata lives on the message. Dropping that source leaves the message and generation_info in the same priority order as before, expressed as a loop rather than a table. Generated-By: PostHog Desktop Task-Id: 5cb23cd6-224c-4d86-bc24-97730d4c2f98
1 parent c636ebd commit 09a8c4e

7 files changed

Lines changed: 181 additions & 15 deletions

File tree

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,5 @@
1+
---
2+
pypi/posthog: patch
3+
---
4+
5+
Only terminal Responses API statuses become `$ai_stop_reason`: a queued or in-progress background run no longer records a lifecycle state as its stop reason, and an incomplete run is named by what cut it short (`incomplete_details.reason`, e.g. `max_output_tokens`). Streaming runs that end incomplete or failed now carry a stop reason too, and the LangChain callback reads stop reasons from `response_metadata` as well, covering Responses API and Anthropic runs that previously recorded none.

‎posthog/ai/langchain/callbacks.py‎

Lines changed: 23 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -48,6 +48,7 @@
4848
_extract_cache_creation_ttl_breakdown,
4949
finalize_ai_content,
5050
get_model_params,
51+
_responses_stop_reason,
5152
with_privacy_mode,
5253
)
5354
from posthog.client import Client
@@ -716,14 +717,11 @@ def _capture_generation(
716717
finalize_ai_content(completions, self._ph_client),
717718
)
718719

719-
# Extract stop reason from generation info
720+
# Extract the stop reason from the generation and its metadata
720721
if output.generations and output.generations[-1]:
721-
last_gen = output.generations[-1][-1]
722-
gen_info = getattr(last_gen, "generation_info", None)
723-
if isinstance(gen_info, dict):
724-
finish_reason = gen_info.get("finish_reason")
725-
if finish_reason is not None:
726-
event_properties["$ai_stop_reason"] = finish_reason
722+
stop_reason = _extract_stop_reason(output.generations[-1][-1])
723+
if stop_reason is not None:
724+
event_properties["$ai_stop_reason"] = stop_reason
727725

728726
_capture_ai_event(
729727
self._ph_client,
@@ -745,6 +743,24 @@ def _log_debug_event(
745743
)
746744

747745

746+
def _extract_stop_reason(generation: Any) -> Optional[str]:
747+
"""
748+
Providers report the stop reason on the message's `response_metadata` or in
749+
`generation_info`, under either spelling. The Responses API reports no
750+
finish reason at all, so a terminal status stands in for one.
751+
"""
752+
metadata = getattr(getattr(generation, "message", None), "response_metadata", None)
753+
info = getattr(generation, "generation_info", None)
754+
755+
for source in (metadata, info):
756+
for key in ("finish_reason", "stop_reason"):
757+
value = source.get(key) if isinstance(source, dict) else None
758+
if value is not None:
759+
return str(value)
760+
761+
return _responses_stop_reason(metadata)
762+
763+
748764
def _extract_raw_response(last_response):
749765
"""Extract the response from the last response of the LLM call."""
750766
# We return the text of the response if not empty

‎posthog/ai/openai/_streaming.py‎

Lines changed: 7 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
from typing import Any, Dict, List, Optional
55

66
from ..types import StreamingEventData, TokenUsage
7-
from ..utils import merge_usage_stats
7+
from ..utils import merge_usage_stats, _responses_stop_reason
88
from .openai_converter import (
99
accumulate_openai_tool_calls,
1010
extract_openai_content_from_chunk,
@@ -38,10 +38,12 @@ def process_chunk(self, chunk: Any) -> None:
3838
if content is not None:
3939
self.output.extend(content)
4040

41-
if getattr(chunk, "type", None) == "response.completed" and response:
42-
status = getattr(response, "status", None)
43-
if status is not None:
44-
self.stop_reason = status
41+
# A stream can end on response.completed, response.incomplete, or
42+
# response.failed; any terminal response names the stop reason.
43+
if response:
44+
stop_reason = _responses_stop_reason(response)
45+
if stop_reason is not None:
46+
self.stop_reason = stop_reason
4547

4648

4749
@dataclass

‎posthog/ai/openai/openai_converter.py‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,7 @@
1717
FormattedTextContent,
1818
TokenUsage,
1919
)
20-
from posthog.ai.utils import serialize_raw_usage
20+
from posthog.ai.utils import _responses_stop_reason, serialize_raw_usage
2121

2222

2323
def _item_attr(item: Any, name: str, default: Any = None) -> Any:
@@ -445,7 +445,7 @@ def extract_openai_stop_reason(response: Any) -> Optional[str]:
445445
return getattr(response.choices[0], "finish_reason", None)
446446
# Responses API
447447
if hasattr(response, "status"):
448-
return getattr(response, "status", None)
448+
return _responses_stop_reason(response)
449449
return None
450450

451451

‎posthog/ai/utils.py‎

Lines changed: 29 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -319,6 +319,35 @@ def format_response(response, provider: str):
319319
return []
320320

321321

322+
# A Responses API run is only finished on these statuses; `queued` and
323+
# `in_progress` are lifecycle states a background run passes through.
324+
_TERMINAL_RESPONSE_STATUSES = frozenset(
325+
{"completed", "failed", "cancelled", "incomplete"}
326+
)
327+
328+
329+
def _read_response_field(source: Any, key: str) -> Any:
330+
return source.get(key) if isinstance(source, dict) else getattr(source, key, None)
331+
332+
333+
def _responses_stop_reason(response: Any) -> Optional[str]:
334+
"""
335+
Map a Responses API outcome to a `$ai_stop_reason`: an incomplete run is
336+
named by what cut it short (`incomplete_details.reason`, e.g.
337+
`max_output_tokens`), the other terminal statuses stand for themselves,
338+
and a non-terminal status yields None. Accepts an SDK response object or
339+
a LangChain `response_metadata` dict.
340+
"""
341+
status = _read_response_field(response, "status")
342+
if not isinstance(status, str) or status not in _TERMINAL_RESPONSE_STATUSES:
343+
return None
344+
details = _read_response_field(response, "incomplete_details")
345+
reason = _read_response_field(details, "reason")
346+
if status == "incomplete" and isinstance(reason, str) and reason:
347+
return reason
348+
return status
349+
350+
322351
def extract_stop_reason(response: Any, provider: str) -> Optional[str]:
323352
"""Extract stop reason from response based on provider."""
324353
if provider == "openai":

‎posthog/test/ai/langchain/test_callbacks.py‎

Lines changed: 55 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2864,3 +2864,58 @@ def test_served_service_tier_merges_into_model_parameters(mock_client):
28642864
props = mock_client.capture.call_args.kwargs["properties"]
28652865
assert props["$ai_model_parameters"]["service_tier"] == "flex"
28662866
assert props["$ai_model_parameters"]["temperature"] == 0.5
2867+
2868+
2869+
@pytest.mark.parametrize(
2870+
"generation_info,response_metadata,expected",
2871+
[
2872+
# generation_info finish_reason keeps priority
2873+
({"finish_reason": "stop"}, {"status": "completed"}, "stop"),
2874+
# Responses API: a terminal status carries no finish_reason at all
2875+
(None, {"status": "completed", "incomplete_details": None}, "completed"),
2876+
# ... an incomplete run is named by what cut it short
2877+
(
2878+
None,
2879+
{
2880+
"status": "incomplete",
2881+
"incomplete_details": {"reason": "max_output_tokens"},
2882+
},
2883+
"max_output_tokens",
2884+
),
2885+
# a queued background run has no stop reason yet
2886+
(None, {"status": "queued", "incomplete_details": None}, None),
2887+
# Anthropic reports through response_metadata.stop_reason
2888+
(None, {"stop_reason": "end_turn"}, "end_turn"),
2889+
],
2890+
)
2891+
def test_stop_reason_resolution(
2892+
mock_client, generation_info, response_metadata, expected
2893+
):
2894+
from langchain_core.outputs import ChatGeneration, LLMResult
2895+
2896+
cb = CallbackHandler(mock_client)
2897+
run_id = uuid.uuid4()
2898+
cb._set_llm_metadata(
2899+
serialized={},
2900+
run_id=run_id,
2901+
messages=[{"role": "user", "content": "test"}],
2902+
metadata={"ls_provider": "openai", "ls_model_name": "gpt-4o"},
2903+
)
2904+
response = LLMResult(
2905+
generations=[
2906+
[
2907+
ChatGeneration(
2908+
message=AIMessage(
2909+
content="Response", response_metadata=response_metadata
2910+
),
2911+
generation_info=generation_info,
2912+
)
2913+
]
2914+
],
2915+
llm_output={},
2916+
)
2917+
2918+
cb._pop_run_and_capture_generation(run_id, None, response)
2919+
2920+
props = mock_client.capture.call_args.kwargs["properties"]
2921+
assert props.get("$ai_stop_reason") == expected

‎posthog/test/ai/openai/test_openai_converter.py‎

Lines changed: 60 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,5 @@
1+
import types
2+
13
import pytest
24

35
try:
@@ -7,7 +9,11 @@
79
except ImportError:
810
OPENAI_AVAILABLE = False
911

10-
from posthog.ai.openai.openai_converter import format_openai_input
12+
from posthog.ai.openai._streaming import _ResponsesStreamState
13+
from posthog.ai.openai.openai_converter import (
14+
extract_openai_stop_reason,
15+
format_openai_input,
16+
)
1117
from posthog.test.ai.utils import make_response_usage
1218

1319
pytestmark = pytest.mark.skipif(not OPENAI_AVAILABLE, reason="openai not available")
@@ -198,3 +204,56 @@ def _chunk(delta_kwargs):
198204

199205
refusal_block = next(b for b in content if b["type"] == "refusal")
200206
assert refusal_block["refusal"] == "I can't help with that"
207+
208+
209+
def _response(status, incomplete_reason=None, **extra):
210+
details = (
211+
types.SimpleNamespace(reason=incomplete_reason) if incomplete_reason else None
212+
)
213+
return types.SimpleNamespace(status=status, incomplete_details=details, **extra)
214+
215+
216+
@pytest.mark.parametrize(
217+
"status,incomplete_reason,expected",
218+
[
219+
# Terminal statuses stand for themselves
220+
("completed", None, "completed"),
221+
("failed", None, "failed"),
222+
("cancelled", None, "cancelled"),
223+
# An incomplete run is named by what cut it short
224+
("incomplete", "max_output_tokens", "max_output_tokens"),
225+
("incomplete", None, "incomplete"),
226+
# Lifecycle states of a background run are not stop reasons
227+
("queued", None, None),
228+
("in_progress", None, None),
229+
],
230+
)
231+
def test_extract_stop_reason_maps_responses_statuses(
232+
status, incomplete_reason, expected
233+
):
234+
assert extract_openai_stop_reason(_response(status, incomplete_reason)) == expected
235+
236+
237+
@pytest.mark.parametrize(
238+
"status,incomplete_reason,expected",
239+
[
240+
("completed", None, "completed"),
241+
("incomplete", "max_output_tokens", "max_output_tokens"),
242+
("failed", None, "failed"),
243+
("in_progress", None, None),
244+
],
245+
)
246+
def test_responses_stream_records_every_terminal_stop_reason(
247+
status, incomplete_reason, expected
248+
):
249+
state = _ResponsesStreamState()
250+
state.process_chunk(
251+
types.SimpleNamespace(
252+
type=f"response.{status}",
253+
response=_response(
254+
status, incomplete_reason, model="gpt-4o", usage=None, output=[]
255+
),
256+
)
257+
)
258+
259+
assert state.stop_reason == expected

0 commit comments

Comments
 (0)