HC 모드 provider final 미확보 시 판독 불가 되묻기 복구

마지막 업데이트 2026-10-01

ba394274 charles-na · 2026-10-01 Fix 8 files +334 −33 PR #1148

Half Cascade(HC) 모드에서 2.5초 안에 OpenAI provider final을 확보하지 못하면, vendor AgentActivity가 턴을 폐기하고 아무 응답 없이 끝났다(이현우 18회기 AI 무응답). 이 PR은 폐기 지점에서 새 세션 이벤트를 emit하고, agent.py가 이를 받아 VP 모드와 같은 "(판독 불가)" 기록 + LLM 되묻기로 이어지게 한다. 리뷰에서 가장 먼저 볼 지점은 두 가지다. ① vendor emit 게이트(agent_activity.py:3774) ② 로컬 commit 전에 도착한 실패를 보류했다가 commit 시 평가하는 경합 처리(agent.py:3499, :3619).

무엇이, 왜 바뀌었나

VP 모드(completion_contract="local_stt")에는 STT final이 오지 않을 때 쓰는 판독 불가 복구가 이미 있었다. HC(vendor_exact)에서는 다음 세 곳에서 막혀 복구가 동작하지 않았다.

이 PR은 ③에 신호를 추가하고, 복구 판정(_stt_no_final_turn_is_eligible)과 복구 실행(_deliver_unreadable_recovery)은 VP 것을 그대로 재사용한다.

데이터 입출력 흐름 — 폐기 턴이 되묻기가 되기까지

1

vendor: 2.5초 안에 provider final 미확보 → 턴 폐기

vendor/…/voice/agent_activity.py:3771

입력: _EndOfTurnInfo(input_turn_id, skip_reply), commit 예약. 출력: 경고 로그 + 새 이벤트

2

vendor: user_input_transcription_failed emit

agent_session.py:2258 events.py:351

페이로드: {input_turn_id, reason="provider_final_not_admitted", created_at}

3

agent.py: vendor 토큰을 로컬 턴으로 매핑하고 보류

agent.py:4179 input_gate.py:155

입력: input_turn_id. 출력: 세션로그 user_stt_provider_final_failed(mapped/superseded), 보류 집합에 turn_id 추가

4

agent.py: 로컬 commit 확인 후 복구 대상인지 판정하고 선점

agent.py:3499 :3619 :3177 :3209

commit 전이면 대기. commit 시점에 다시 평가한다. 복구 대상이 아니면 보류를 폐기하고 끝낸다.

5

공용 복구: "(판독 불가)" 로그 + 되묻기 응답 생성

agent.py:3363 LiveKit data → 브라우저 SessionLog

출력: user_stt_provider_final_fallback, user_stt_segment "(판독 불가)", clear_user_turn → interrupt(force) → generate_reply(되묻기 지시문)

시간순 인과 체인

# 아동 발화 종료 VAD silence ─┬─▶ agent.py endpointing 타이머 ──▶ turn_committed_at 기록 (:3619) └─▶ vendor EOT ──▶ commit 예약 ──▶ provider final 대기(≤2.5초) │ final 미확보 / 빈 final / commit 실패 ▼ user_input_transcription_failed(input_turn_id) ▼ _on_user_input_transcription_failed (:4179) vendor_turn_epoch → turn_id ✗ 매핑 실패 → 진단 로그만 is_latest_vendor_turn ✗ superseded(resume된 이전 토큰) → 진단 로그만 pending.add(turn_id) ▼ _maybe_schedule_provider_final_fallback (:3499) committed 전? ── 예 ──▶ 보류 유지 → commit 시점(:3619)에 다시 호출 eligible? ── 아니오 ─▶ 보류 폐기 (새 발화·다른 응답·듣기 OFF·취소 세대·종료) _claim_unreadable_recovery_turn # 동기 선점, await 전 ▼ _deliver_unreadable_recovery (:3363, VP와 공유) → "(판독 불가)" → clear_user_turn → interrupt(force) → generate_reply ▼ Realtime LLM이 되묻기 문장 생성 → Typecast TTS 재생

코드로 보는 핵심 지점

1. vendor 폐기 지점의 emit 게이트

apps/livekit-agent/vendor/livekit-agents/livekit/agents/voice/agent_activity.py @ _user_turn_completed_impl logger.warning("skipping Half Cascade turn without exact provider final", ...) self._cancel_preemptive_generation_if_current(turn_preemptive_generation) if ( info.input_turn_id and not info.skip_reply and not self._session._closing and realtime_audio_commit.session is self._rt_session ): self._session._user_input_transcription_failed( UserInputTranscriptionFailedEvent(input_turn_id=info.input_turn_id) ) return

reservation.invalidated는 쓰지 않았다. 정상 실패·clear·종료가 모두 이 플래그를 True로 만들어 구분용으로 쓸 수 없기 때문이다. vendor에서는 세션 종료 flush·종료 중·세션 교체만 거르고, 수동 인터럽트처럼 여기서 구분할 수 없는 경우는 agent.py의 eligibility 검사가 막는다.

2. 실패 이벤트 수신과 resume된 이전 토큰 걸러내기

apps/livekit-agent/agent.py:4179 @ _attach_event_logger @session.on("user_input_transcription_failed") def _on_user_input_transcription_failed(ev) -> None: turn_id = turn_guard.vendor_turn_epoch(input_turn_id) ... superseded = turn_id is not None and not turn_guard.is_latest_vendor_turn(input_turn_id) _publish({"type": "user_stt_provider_final_failed", "mapped": ..., "superseded": ...}) if turn_id is None or superseded or turn_id in stt_fallback_turn_ids: return pending_provider_final_failure_turn_ids.add(turn_id) _maybe_schedule_provider_final_fallback(turn_id) apps/livekit-agent/input_gate.py:155 def is_latest_vendor_turn(self, input_turn_id) -> bool: epoch_id = self.vendor_turn_epoch(input_turn_id) return epoch_id is not None and self._latest_vendor_turn_by_epoch.get(epoch_id) == input_turn_id

리뷰 1라운드 P1을 고친 부분이다. 아동이 말을 멈췄다가 바로 이어 말하면 같은 epoch에 input-1·input-2가 붙는다. 이때 input-1의 실패가 input-2의 정상 발화를 지울 수 있었다. 지금은 최신 토큰이 아니면 superseded로 기록만 한다. _resume_endpointing_turn(:3117)도 보류해 둔 실패를 버린다.

3. commit 전 도착분 보류 → commit 시점 평가

apps/livekit-agent/agent.py:3499 @ _maybe_schedule_provider_final_fallback if (turn_id not in pending_provider_final_failure_turn_ids or "turn_committed_at" not in turn_times.get(turn_id, {})): return False # commit 전: 보류 유지 pending_provider_final_failure_turn_ids.discard(turn_id) if turn_id in unreadable_recovery_tasks or not _stt_no_final_turn_is_eligible(...): return False _claim_unreadable_recovery_turn(turn_id, "provider_final_fallback") unreadable_recovery_tasks[turn_id] = loop.create_task(_run_provider_final_fallback(turn_id)) apps/livekit-agent/agent.py:3617 @ _run_terminalization_after_delay if not _maybe_schedule_soniox_blank_fallback(turn_id): if not ( _maybe_schedule_soniox_blank_fallback(turn_id) or _maybe_schedule_provider_final_fallback(turn_id) ): _arm_stt_no_final_watchdog(turn_id)

commit 실패나 빈 final은 대기 없이 즉시 실패하므로, 실패 이벤트가 agent.py의 endpointing commit보다 먼저 올 수 있다. 그래서 보류해 두었다가 commit 시점에 다시 평가한다.

4. 공용 선점 헬퍼와 task 저장소 공유 (구조 변경)

apps/livekit-agent/agent.py:3209 def _claim_unreadable_recovery_turn(turn_id, initiator) -> None: stt_fallback_turn_ids.add(turn_id) (terminalized 대기열에서 제거) turn_guard.mark_turn_dropped_by_terminalization(turn_id, initiator) turn_guard.record_cancellation(initiator=initiator, cleared_user_turn_id=turn_id) apps/livekit-agent/agent.py (전역 이름 변경) soniox_blank_fallback_tasks: dict[str, asyncio.Task] = {} unreadable_recovery_tasks: dict[str, asyncio.Task] = {}

VP no-final·Soniox blank 경로에 있던 같은 내용의 선점 블록 두 개를 헬퍼 하나로 묶었다. 사이에 await가 없어서 순서가 바뀌어도 결과는 같다. dict 이름 변경은 이벤트 계약 변경이 아니다. Soniox blank와 HC 복구 task를 같은 저장소에 넣어서, 중복 방지(:3450·:3516)와 새 발화·종료 시 일괄 취소(:3580)를 두 경로에 함께 적용하기 위한 것이다.

레이어별 변경 요약 · 코드 변경점 위치

레이어파일 (위치)핵심 변경
vendor 계약vendor/…/voice/events.py:351UserInputTranscriptionFailedEvent 추가, EventTypes·AgentEvent 등록
vendor 계약vendor/…/voice/__init__.py이벤트 export
vendor 세션vendor/…/voice/agent_session.py:2258emit 헬퍼 _user_input_transcription_failed
vendor 런타임vendor/…/voice/agent_activity.py:3774폐기 지점 emit + 게이트 (+12)
입력 게이트input_gate.py:155, agent.py:2132is_latest_vendor_turn + 래퍼
agent 상태agent.py:2577, :2579unreadable_recovery_tasks(이름 변경), pending_provider_final_failure_turn_ids(신규)
agent 수신agent.py:4179실패 이벤트 핸들러: 매핑·superseded 판정·보류
agent 스케줄agent.py:3485, :3499, :3619HC 복구 스케줄러·실행 task, commit 시점 연결
agent 공용agent.py:3209_claim_unreadable_recovery_turn 추출 (no-final·Soniox·HC 공유)
agent 정리agent.py:3071, :3117, :3573턴 정리·resume·새 발화/종료 시 보류 실패 폐기
테스트tests/test_realtime_livekit_vad_identity.py폐기 시 이벤트 1회, 세션 교체 시 0회 (기존 2개 테스트 보강)
테스트tests/test_turn_terminalization.pyHC 복구 1회(중복 포함), commit 전 도착, 억제 3종, resume 전·후 superseded (+155)

리뷰 관전 포인트

정책 미결정

빈 provider final도 즉시 되묻기: deadline 만료·빈 transcript·admission 실패가 모두 같은 provider_final_not_admitted로 나간다. 그래서 기침·잡음에도 AI가 다시 말해 달라고 할 수 있다. VP의 Soniox blank는 soniox_blank_policy로 켜고 끌 수 있다. 사유를 timeout/blank로 나누고 blank를 정책값으로 켜고 끌지 결정이 필요하다.

구조

vendor 수정: LiveKit agents를 업그레이드할 때 이 패치를 다시 적용하고 검증해야 한다. 로컬 venv는 vendor를 non-editable로 설치하므로, vendor를 수정하면 uv sync --frozen --reinstall-package livekit-agents를 해야 테스트에 반영된다.

동작 확인

응답 지연: 되묻기 문장을 Realtime LLM이 만들므로, 폐기 시점(EOT + 2.5초) 뒤에 응답 한 번을 더 기다린다. 실기기에서 체감 지연과, 세션로그에 failed → fallback → "(판독 불가)" 순서가 남는지 확인이 필요하다.

테스트 공백

다음 경우를 검증하는 테스트가 없다.

  • resume 후 input-2가 실패했을 때 정상 복구되는 positive 케이스
  • vendor의 skip_reply/_closing 가드
  • 실패 사이에 수동 인터럽트가 끼는 경우
  • commit 전 실패 뒤 늦은 final이 오는 경우
P3
  • vendor의 _closing은 close가 시작된 뒤에야 True가 된다. _is_closing()을 쓰면 더 일찍 막을 수 있다.
  • 잠깐 부적격이어서 보류를 폐기하면 그 턴의 복구 기회가 사라진다. 기존 VP와 같은 동작이다.
  • 로컬 outcome이 dropped인데 실패 이벤트가 그보다 늦게 오면, 보류 값이 다음 발화나 종료 때에야 정리된다.
영향 범위

테스트: 전체 491개가 통과했다. 새 테스트는 구현 전 또는 수정 전에 실패하는 것을 확인했다.

리뷰: guardian-gate 2라운드에서 P0는 없었고, P1 1건은 수정 후 해소됐다.

영향 없음: 정상 HC 턴, VP 경로, 웹 클라이언트. 웹은 세션로그에 이벤트를 표시만 한다.

관련 문서