Skip to content

[Feat] P4 — Layer C 핵심 로직 구현 (Leiden community detection + map-reduce 요약 합성) #59

Description

@TaskerJang

요약

Layer C 의 진입점 / 라우팅 분기 / stub 안내는 PR #51에서 이미 구현 완료 (retrieval/community_summary.py). 본 이슈는 stub 안내 응답을 진짜 답변으로 대체 — Leiden community detection + community summary 사전 생성 + map-reduce 합성.

Hybrid 로드맵 P4 📦 — multi-doc 영역 회복.

동기 — W6 측정에서 진단된 영향

Pattern 영향 건수 증상
multi_doc_trend 10 (routing 100%) 라우터가 Layer C 로 옳게 분기 → stub 응답 반환 → Correctness 1.0
summary 4 corpus-wide 요약 필요 → BM25 chunk 검색 부적합
그 외 글로벌 질의 2 "전체 트렌드" / "주요 흐름" 패턴

총 16건의 부진이 Layer C 진짜 구현으로 회복 가능.

라우터는 정상 작동 (multi_doc_trend Routing Accuracy 100%) — 답변 단계만 박으면 됨.

현재 상태 — PR #51 박힘

# retrieval/community_summary.py — 이미 구현된 부분
@track
def community_summary(question: str) -> CommunitySummaryResult:
    # 진입점 ✅
    # @track Opik trace ✅
    # CommunitySummaryResult 시그니처 박제 ✅
    # is_stub=True 정직한 안내 응답 ✅

    return CommunitySummaryResult(
        question=question,
        answer=_STUB_ANSWER_TEMPLATE.format(question=question),
        is_stub=True,  # ← 본 이슈로 False 전환
    )

변경 사항 (예정)

1. kg/community_detection.py (신규)

Leiden 알고리즘 박힌 community detection:

def detect_communities(
    *,
    resolution: float = 1.0,
    seed: int = 42,
) -> List[Community]:
    """Entity 노드 기반 Leiden community 추출.

    Returns:
        List[Community] — 각 community 는 entity_ids + level + score 박힘.
    """
    # 1. Neo4j GDS 프로젝션 (entity 노드 + 관계)
    # 2. Leiden 알고리즘 실행
    # 3. Community 메타데이터 → Neo4j 노드 박기 (Community 라벨)
    ...

후보 라이브러리:

  • Neo4j GDS (gds.leiden.write) — Neo4j 내장, 가장 깔끔
  • networkx + leidenalg — Python 단독 가능

2. kg/community_summary.py (신규)

각 community 의 LLM 요약 사전 생성 (indexing-time):

def build_community_summaries(
    communities: List[Community],
    *,
    llm_config: LLMConfig,
) -> List[CommunityReport]:
    """각 community → LLM 요약 생성 → Neo4j 노드 박기 (CommunityReport 라벨).

    각 report 박힘:
    - entities: 해당 community 의 entity 목록
    - relationships: community 내부 관계
    - core_chunks: 핵심 chunk 텍스트 3~5개
    - summary: LLM 생성 요약 (한국어, ~500 토큰)
    """
    ...

3. retrieval/community_summary.py (수정 — stub → 실제 구현)

@track
def community_summary(question: str) -> CommunitySummaryResult:
    # 1. Map step — 모든 CommunityReport 에 대해 question 관련성 점수
    relevant_reports = _map_step(question, llm_config)

    # 2. Reduce step — top-K report 기반 최종 답변 합성
    answer = _reduce_step(question, relevant_reports, llm_config)

    return CommunitySummaryResult(
        question=question,
        answer=answer,
        is_stub=False,  # ← 진짜 구현
        community_count=len(relevant_reports),
        top_communities=[r.community_id for r in relevant_reports[:3]],
    )

4. scripts/build_community_index.py (신규)

Indexing-time community 빌드 진입점:

uv run python scripts/build_community_index.py \
  --algorithm leiden \
  --resolution 1.0 \
  --llm-model "moonshotai/kimi-k2.5"

5. retrieval/prompts/community_summary_v2.md (수정)

기존 placeholder 박힌 거 → 실제 map-reduce 프롬프트로 교체.

6. docs/adr/0007-community-detection.md (신규)

Leiden vs Louvain vs Label Propagation 의사결정 ADR.

DoD

  • scripts/build_community_index.py 박혀서 community 빌드 성공
    • Community 수 5~15개 박혀있어야 (8 doc 기준 적정)
    • 각 community 당 entity 5~30개 박혀있어야
  • CommunityReport 노드 박혀서 LLM 요약 박혀있어야
  • community_summary(question)is_stub=False, 실제 답변 박힘
  • R3 케이스 (W4 평가셋) 다시 박아서 stub 응답 아닌 진짜 corpus 요약 답변 박혀있어야

검증 방법

# 1. P1 + P2 적용된 그래프 기반 community 빌드
uv run python scripts/build_community_index.py --algorithm leiden

# 2. Community 노드 audit
uv run python scripts/audit_communities.py
# Expected: Community 5~15개, 각 5~30 entities, summary 박혀있음

# 3. W4 R3 케이스 재실행
uv run python -m scripts.run_w4_eval --case R3
# Expected: is_stub=False, 답변에 corpus-wide 트렌드 박혀있음

# 4. 80 QA 재측정 (multi_doc_trend 10건 + summary 4건 회복 확인)
uv run python scripts/run_qa_eval.py --qa-set graphrag --tag layer_c_v2

예상 효과

Pattern Before (W6, stub) After (P4, 진짜) 회복
multi_doc_trend Correctness 1.00 3.0+ 10건
summary Correctness 1.00 3.0+ 4건
Faithful (전체) 5.0% +5%p 추가 6건

Trade-offs

항목 비용
Indexing-time 빌드 시간 8 doc / 515 community → 약 1030분
LLM 호출 (community 요약) 5~15회 (community 당 1회)
Neo4j 저장소 CommunityReport 노드 추가

indexing-time 비용이라 query-time 영향 없음 (빠른 응답 유지).

의존 / 후속

레퍼런스

우선순위

📦 P4 — P1+P2 적용 후 진행. 2~3일 작업.


Refs #20, #51, #56, #57, W6 measurement results

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions