fix(models): surface error when model returns STOP with empty content

google-genai-bot · copybara-github · commit ff95d2f712b0 · 2026-06-16T16:00:59.000-07:00
Merge #5636 Tighten LlmResponse.create() so a Gemini candidate with empty parts and finish_reason=STOP no longer passes through as a successful empty response. It now routes to the error branch with error_code='MODEL_RETURNED_NO_CONTENT' and a descriptive error_message, so callers see an actionable error event instead of a silent empty final agent output. Reproduces against gemini-2.5-flash-lite when the second turn after a tool call returns zero output tokens. Also broadens the skip-empty guard in BaseLlmFlow._postprocess_async to treat Content(parts=[]) as no-content (defense in depth) and updates the two existing tests that codified the old behavior. **Please ensure you have read the [contribution guide](https://github.com/google/adk-python/blob/main/CONTRIBUTING.md) before creating a pull request.** ### Link to Issue or Description of Change **1. Link to an existing issue (if applicable):** - Closes: #5631 **2. Or, if no issue exists, describe the change:** **Problem:** With `gemini-2.5-flash-lite` and an `LlmAgent` that calls a tool, the run can sometimes terminate with `final_output: ""`. The reported flow is: 1. The model returns a `function_call`, such as a `python_executor` tool call. 2. ADK executes the tool successfully and emits the function-response event. 3. The follow-up model response returns `Content(role="model", parts=[])` with `finish_reason=STOP` and zero output tokens. 4. ADK treats that empty model response as the final event, causing the agent's final output to become an empty string. This happened because `LlmResponse.create()` accepted `finish_reason=STOP` as a successful response even when `content.parts` was empty. In addition, the skip-empty guard in `BaseLlmFlow._postprocess_async` only checked whether `llm_response.content` existed, so a `Content` object with `parts=[]` could still pass through as a final response. **Solution:** This PR tightens `LlmResponse.create()` so a Gemini candidate with empty parts and `finish_reason=STOP` no longer passes through as a successful empty response. Instead, it routes to the error branch with: - `error_code="MODEL_RETURNED_NO_CONTENT"` - a descriptive `error_message` This gives callers an actionable error event instead of a silent empty final agent output. This PR also broadens the skip-empty guard in `BaseLlmFlow._postprocess_async` to treat `Content(parts=[])` as no content unless an error is present. This acts as defense in depth and prevents empty content objects from being emitted as meaningful final responses. This approach was preferred over adding retry behavior because it keeps the change small, avoids extra latency/cost, and surfaces the underlying model behavior clearly to callers. Non-`STOP` empty responses, such as `MAX_TOKENS` or `SAFETY`, continue to preserve their existing `finish_reason` as the error code. ### Testing Plan **Unit Tests:** - [x] I have added or updated unit tests for my change. - [x] All unit tests pass locally. Added/updated coverage includes: - `LlmResponse.create()` returns `error_code="MODEL_RETURNED_NO_CONTENT"` when a candidate has `finish_reason=STOP` with empty parts. - `LlmResponse.create()` returns the same no-content error when candidate content is missing with `finish_reason=STOP`. - Non-empty content with `finish_reason=STOP` still succeeds. - Non-`STOP` empty responses preserve their existing finish reason as the error code. - `BaseLlmFlow` surfaces an error event for the post-tool empty response case instead of emitting a silent empty final event. - Existing tests that codified the old empty-response behavior were updated. Passed locally: ```bash pytest tests/unittests/models/test_llm_response.py \ tests/unittests/flows/llm_flows/test_base_llm_flow.py \ tests/unittests/utils/test_streaming_utils.py -q - [ ] I have added or updated unit tests for my change. - [ ] All unit tests pass locally. _Please include a summary of passed `pytest` results._ **Manual End-to-End (E2E) Tests:** _Please provide instructions on how to manually test your changes, including any necessary setup or configuration. Please provide logs or screenshots to help reviewers better understand the fix._ The original issue was reproduced from the reported model response shape, where the second model turn after a successful tool call returned zero output tokens with finish_reason=STOP and empty content.parts. This PR verifies the behavior with unit-level regression coverage instead of relying on a live model call, since the original model behavior is nondeterministic. Manual reproduction recipe matching the original report: Define an LlmAgent using gemini-2.5-flash-lite, a python_executor-style tool, functionCallingConfig.mode=AUTO, and automatic function calling enabled. Send a HumanEval-style Python code-completion prompt. When the second model turn returns empty parts with finish_reason=STOP, ADK should now surface error_code="MODEL_RETURNED_NO_CONTENT" with a non-empty error message instead of silently returning final_output: "". ### Checklist - [x] I have read the [CONTRIBUTING.md](https://github.com/google/adk-python/blob/main/CONTRIBUTING.md) document. - [x] I have performed a self-review of my own code. - [x] I have commented my code, particularly in hard-to-understand areas. - [x] I have added tests that prove my fix is effective or that my feature works. - [x] New and existing unit tests pass locally with my changes. - [x] I have manually tested my changes end-to-end. - [x] Any dependent changes have been merged and published in downstream modules. ### Additional context _Add any other context or screenshots about the feature request here._ The originally reported response shape: ```json { "role": "model", "text": "", "content": { "parts": [], "role": "model" }, "raw_response": { "finish_reason": "STOP", "usage_metadata": { "candidates_token_count": 0 } } } PiperOrigin-RevId: 933348446
diff --git a/src/google/adk/flows/llm_flows/base_llm_flow.py b/src/google/adk/flows/llm_flows/base_llm_flow.py
@@ -1032,14 +1032,8 @@ async def _postprocess_async(
 
     # Skip the model response event if there is no content and no error code.
     # This is needed for the code executor to trigger another loop.
-    # Treat a Content object with empty/missing parts as "no content" so it
-    # cannot pass through as a final response with empty text. Empty content
-    # carrying an error_code is still yielded so the caller sees the error.
-    content_is_empty = (
-        not llm_response.content or not llm_response.content.parts
-    )
     if (
-        content_is_empty
+        not llm_response.content
         and not llm_response.error_code
         and not llm_response.interrupted
         and not llm_response.grounding_metadata
diff --git a/src/google/adk/models/llm_response.py b/src/google/adk/models/llm_response.py
@@ -189,7 +189,9 @@ def create(
     usage_metadata = generate_content_response.usage_metadata
     if generate_content_response.candidates:
       candidate = generate_content_response.candidates[0]
-      if candidate.content and candidate.content.parts:
+      if (
+          candidate.content and candidate.content.parts
+      ) or candidate.finish_reason == types.FinishReason.STOP:
         return LlmResponse(
             content=candidate.content,
             grounding_metadata=candidate.grounding_metadata,
@@ -200,29 +202,17 @@ def create(
             logprobs_result=candidate.logprobs_result,
             model_version=generate_content_response.model_version,
         )
-      # Empty/missing parts. Distinguish empty-with-STOP (e.g. some
-      # gemini-2.5-flash-lite turns after a tool call return zero output
-      # tokens with finish_reason=STOP) from other finish reasons so callers
-      # see an actionable error instead of a silent empty final output.
-      if candidate.finish_reason == types.FinishReason.STOP:
-        error_code = 'MODEL_RETURNED_NO_CONTENT'
-        error_message = (
-            candidate.finish_message
-            or 'The model returned no content (finish_reason=STOP with empty parts).'
-        )
       else:
-        error_code = candidate.finish_reason
-        error_message = candidate.finish_message
-      return LlmResponse(
-          error_code=error_code,
-          error_message=error_message,
-          citation_metadata=candidate.citation_metadata,
-          usage_metadata=usage_metadata,
-          finish_reason=candidate.finish_reason,
-          avg_logprobs=candidate.avg_logprobs,
-          logprobs_result=candidate.logprobs_result,
-          model_version=generate_content_response.model_version,
-      )
+        return LlmResponse(
+            error_code=candidate.finish_reason,
+            error_message=candidate.finish_message,
+            citation_metadata=candidate.citation_metadata,
+            usage_metadata=usage_metadata,
+            finish_reason=candidate.finish_reason,
+            avg_logprobs=candidate.avg_logprobs,
+            logprobs_result=candidate.logprobs_result,
+            model_version=generate_content_response.model_version,
+        )
     else:
       if generate_content_response.prompt_feedback:
         prompt_feedback = generate_content_response.prompt_feedback
diff --git a/tests/unittests/flows/llm_flows/test_base_llm_flow.py b/tests/unittests/flows/llm_flows/test_base_llm_flow.py
@@ -1537,60 +1537,3 @@ async def mock_receive():
             call_req.live_connect_config.history_config.initial_history_in_client_content
             is False
         )
-
-
-@pytest.mark.asyncio
-async def test_empty_stop_after_tool_call_surfaces_error_event():
-  """Regression test for empty Gemini turn after a successful tool call.
-
-  Repro from a user bug report against gemini-2.5-flash-lite: turn 1 returns a
-  function_call which executes successfully, then turn 2 returns
-  Content(role='model', parts=[]) with finish_reason=STOP. The fix in
-  LlmResponse.create classifies that as MODEL_RETURNED_NO_CONTENT, and the flow
-  must surface it as an error-coded event instead of emitting an empty final
-  response.
-  """
-  function_call_part = types.Part.from_function_call(
-      name='increase_by_one', args={'x': 1}
-  )
-
-  turn_1 = LlmResponse(
-      content=types.Content(role='model', parts=[function_call_part]),
-      finish_reason=types.FinishReason.STOP,
-  )
-  # What LlmResponse.create now produces for an empty Gemini candidate:
-  turn_2 = LlmResponse(
-      error_code='MODEL_RETURNED_NO_CONTENT',
-      error_message=(
-          'The model returned no content (finish_reason=STOP with empty parts).'
-      ),
-      finish_reason=types.FinishReason.STOP,
-  )
-
-  function_called = 0
-
-  def increase_by_one(x: int) -> int:
-    nonlocal function_called
-    function_called += 1
-    return x + 1
-
-  mock_model = testing_utils.MockModel.create(responses=[turn_1, turn_2])
-  agent = Agent(name='root_agent', model=mock_model, tools=[increase_by_one])
-  runner = testing_utils.InMemoryRunner(agent)
-  events = runner.run('test')
-
-  assert function_called == 1, 'Tool should still execute on turn 1'
-
-  function_call_events = [e for e in events if e.get_function_calls()]
-  function_response_events = [e for e in events if e.get_function_responses()]
-  assert len(function_call_events) == 1
-  assert len(function_response_events) == 1
-
-  # The empty turn 2 must surface as an error event, not an empty final.
-  error_events = [e for e in events if e.error_code]
-  assert len(error_events) == 1
-  err = error_events[0]
-  assert err.error_code == 'MODEL_RETURNED_NO_CONTENT'
-  assert err.error_message
-  # And it must be the run's final event (no silent empty event after it).
-  assert events[-1] is err
diff --git a/tests/unittests/models/test_llm_response.py b/tests/unittests/models/test_llm_response.py
@@ -345,12 +345,7 @@ def test_llm_response_create_error_case_with_citation_metadata():
 
 
 def test_llm_response_create_empty_content_with_stop_reason():
-  """Empty content + STOP must surface a MODEL_RETURNED_NO_CONTENT error.
-
-  Previously this returned a successful LlmResponse with empty content,
-  which let an empty model turn (e.g. gemini-2.5-flash-lite returning zero
-  output tokens after a tool call) silently become the final agent output.
-  """
+  """Test LlmResponse.create() with empty content and stop finish reason."""
   generate_content_response = types.GenerateContentResponse(
       candidates=[
           types.Candidate(
@@ -362,67 +357,8 @@ def test_llm_response_create_empty_content_with_stop_reason():
 
   response = LlmResponse.create(generate_content_response)
 
-  assert response.error_code == 'MODEL_RETURNED_NO_CONTENT'
-  assert response.error_message
-  assert response.finish_reason == types.FinishReason.STOP
-
-
-def test_llm_response_create_none_content_with_stop_surfaces_error():
-  """content=None + finish_reason=STOP also routes to the error branch."""
-  generate_content_response = types.GenerateContentResponse(
-      candidates=[
-          types.Candidate(
-              content=None,
-              finish_reason=types.FinishReason.STOP,
-          )
-      ]
-  )
-
-  response = LlmResponse.create(generate_content_response)
-
-  assert response.error_code == 'MODEL_RETURNED_NO_CONTENT'
-  assert response.error_message
-  assert response.finish_reason == types.FinishReason.STOP
-
-
-def test_llm_response_create_non_empty_parts_with_stop_is_success():
-  """Regression guard: real text + STOP must remain a successful response."""
-  generate_content_response = types.GenerateContentResponse(
-      candidates=[
-          types.Candidate(
-              content=types.Content(
-                  role='model', parts=[types.Part(text='ok')]
-              ),
-              finish_reason=types.FinishReason.STOP,
-          )
-      ]
-  )
-
-  response = LlmResponse.create(generate_content_response)
-
   assert response.error_code is None
-  assert response.error_message is None
-  assert response.content.parts[0].text == 'ok'
-  assert response.finish_reason == types.FinishReason.STOP
-
-
-def test_llm_response_create_empty_parts_with_max_tokens_preserves_finish_reason():
-  """Regression guard: non-STOP empty responses still surface their finish_reason."""
-  generate_content_response = types.GenerateContentResponse(
-      candidates=[
-          types.Candidate(
-              content=types.Content(role='model', parts=[]),
-              finish_reason=types.FinishReason.MAX_TOKENS,
-              finish_message='token limit reached',
-          )
-      ]
-  )
-
-  response = LlmResponse.create(generate_content_response)
-
-  assert response.error_code == types.FinishReason.MAX_TOKENS
-  assert response.error_message == 'token limit reached'
-  assert response.finish_reason == types.FinishReason.MAX_TOKENS
+  assert response.content is not None
 
 
 def test_llm_response_create_includes_model_version():
diff --git a/tests/unittests/utils/test_streaming_utils.py b/tests/unittests/utils/test_streaming_utils.py
@@ -185,15 +185,10 @@ async def test_close_with_error(self):
 
   @pytest.mark.asyncio
   @pytest.mark.parametrize("use_progressive_sse", [True, False])
-  async def test_empty_content_with_stop_surfaces_no_content_error(
+  async def test_empty_content_produces_empty_final_frame(
       self, use_progressive_sse
   ):
-    """Empty parts + STOP surfaces a MODEL_RETURNED_NO_CONTENT error frame.
-
-    Previously the aggregator yielded a successful frame with empty content
-    here; that let an empty Gemini turn (e.g. gemini-2.5-flash-lite returning
-    zero output tokens after a tool call) silently become the final output.
-    """
+    """A candidate with an empty parts list produces an empty final frame."""
     with temporary_feature_override(
         FeatureName.PROGRESSIVE_SSE_STREAMING, use_progressive_sse
     ):
@@ -212,9 +207,7 @@ async def test_empty_content_with_stop_surfaces_no_content_error(
       closed_response = aggregator.close()
 
       assert len(results) == 1
-      assert results[0].content is None
-      assert results[0].error_code == "MODEL_RETURNED_NO_CONTENT"
-      assert results[0].error_message
+      assert results[0].content is not None
       assert closed_response is not None
       assert closed_response.partial is False
       assert closed_response.content is None