Work/FOOTNOTE
Agentic AI · RAG

FOOTNOTE

Agentic RAG with Sentence-Level Citations

文単位の出典を示すエージェント型RAG

FOOTNOTE screenshot 1FOOTNOTE screenshot 2FOOTNOTE screenshot 3FOOTNOTE screenshot 4

01 / 04
01

Overview

概要

An AI research agent for your documents. It decides what to search for, searches again when the results are thin, and every claim in its answer links to the exact sentence it came from.

社内文書を調べるAIリサーチエージェント。何を検索するかを自分で決め、結果が足りなければ検索し直し、回答のすべての主張に根拠となる一文へのリンクを付けます。

The usual 'chat with your documents' build searches once with the user's words and hopes. A two-part question gets one search and half an answer. footnote replaces that fixed pipeline with an agent that works like a researcher, built around one rule: an answer you cannot check is not an answer.

よくある「文書と対話する」仕組みは、ユーザーの言葉で一度だけ検索して、うまくいくことを祈るだけです。2つの問いを含む質問でも検索は1回で、答えは半分しか返りません。footnoteはこの固定的な流れを、研究者のように動くエージェントに置き換えました。原則は一つ、「検証できない答えは答えではない」です。

02

How it was built

開発

Claude is given two tools, search_handbook and read_page, and loops until it can answer. Search results come back as sentence-level blocks, so every citation is validated by the API rather than typed by the model. Retrieval is a hand-written BM25 plus local bge-small vectors, merged with weighted reciprocal rank fusion and tuned on a 59-question eval set. The demo corpus is the public-domain 18F staff handbook, 232 pages.

Claudeに search_handbook と read_page の2つのツールを与え、答えられるまでループさせます。検索結果は文単位のブロックで返すため、出典はモデルが書いた文字列ではなくAPIが検証したものになります。検索は自作のBM25とローカルのbge-smallベクトルを重み付きRRFで統合し、59問の評価セットで調整しました。デモ用の文書は、パブリックドメインの18F職員ハンドブック(232ページ)です。

03

What it does

機能
01
Plans its own research
調査を自分で計画

A two-part question gets two searches, worded the way the handbook words things. Each step shows live in a side panel.

2つの問いを含む質問には2回検索し、文書側の言い回しで検索語を選びます。各ステップはサイドパネルにリアルタイムで表示。

02
Sentence-level citations
文単位の出典

Click a citation chip and the source passage scrolls into view with the exact sentence highlighted.

出典チップをクリックすると、元の段落が表示され、該当する一文がハイライトされます。

03
Says when it doesn't know
わからない時はそう答える

If nothing relevant turns up after retrying, it says the documents don't cover it instead of guessing.

検索し直しても見つからなければ、推測せずに「文書に記載がない」と答えます。

04
Swappable search
検索方式の切り替え

Run the agent on hybrid, keyword-only or vector-only search, with each passage's rank in both shown.

ハイブリッド・キーワードのみ・ベクトルのみを切り替え可能。各段落の両方式での順位も表示。

04

Stack

技術構成

Python with FastAPI and a one-file HTML UI, Claude through the Anthropic SDK with no framework, and a CLI that prints each tool call. The screenshots come from a keyless demo mode where the model's turns are scripted and the searches and citations run for real.

Python・FastAPI・単一ファイルのHTML UI。フレームワークを使わずAnthropic SDKでClaudeを呼び出し、ツール呼び出しを表示するCLIも用意。スクリーンショットはAPIキー不要のデモモードで撮影しており、モデルの発話は台本ですが、検索と出典の解決は実際に動いています。

Retrieval is the easy half. The hard half is making every sentence of the answer checkable.

検索は簡単な半分。難しいのは、回答のすべての文を検証できるようにすることです。

Next case study 次の実績
FRONTDESK
A Support Agent That Does the Work, With a Human on Refunds
→