Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering

This repository accompanies the FocusGraph paper. FocusGraph selects question-relevant video clips from scene graphs, then applies patchwise sparse-flow retention (PSFR) to the original source video and sends a small set of timestamped frames to a vision-language model.

source video -> 90 clip graphs -> frozen ModernBERT vectors -> GeLM selector
                                                    | 90 clip scores
                                                    v
source video -> continuous-video PSFR cues -> retrieved-frame union -> 8 frames -> QA

The release includes code for scene graph construction, graph embeddings, selector training, frame selection, and question answering on FindingDory and HourVideo. The examples below walk through these stages, from preparing source videos to evaluating answers.

Installation

Run the following commands from this directory using Python 3.11 on Linux. OpenCV and NumPy are pinned to the versions used to compute the HourVideo PSFR features. PSFR runs on the CPU and does not need a GPU.

python3.11 -m venv .venv
.venv/bin/python -m pip install -e '.[test]'
.venv/bin/python -m pytest -q
.venv/bin/focusgraph --help

The package imports its own src/focusgraph code and does not depend on a parent research checkout. After installation, activate the environment with source .venv/bin/activate to use the commands below.

For local Qwen QA, install a CUDA build of PyTorch suitable for your GPU, then run pip install -e '.[qa]'. The [qa,train] extras can be installed together when running the complete selector-to-QA pipeline.

Training requires more resources. The HourVideo-disjoint GeLM experiment used four A100 80 GB GPUs, a global batch size of 64, and a context length of 8192 tokens. The FindingDory checkpoint came from a separate eight-GPU run.

Graph construction uses a separate SG-Ego checkout at commit ecf24c3776b63ad1d9e49272ea090df1d596afdd. Its source is not bundled because the checkout used for this release has no license file. Follow the upstream documentation to install its model dependencies, including CUDA PyTorch, TorchVision, Decord, Transformers, and Hugging Face Hub.

The three graph stages are available as focusgraph-c1, focusgraph-c2, and focusgraph-c3; use --help to see their options. Stage 1 uses Qwen3.5-9B revision c202236235762e1c871ad0ccb60c8ee5ba337b9a. Stages 2 and 3 download the detector/mapper and SAM2/DINOv2 revisions pinned in their source.

Data and model weights

The code repository is accompanied by the precomputed data archives below. Obtain source videos and official benchmark annotations from their original sources. Frozen graph embeddings and model weights must be obtained or generated separately.

Precomputed data downloads

Download the archives from the TO DO: add link. The five archives total approximately 503 MB.

Archive Dataset and split Contents Size
findingdory_graphs.tar.gz FindingDory, full validation split; 100 episodes 90 scene-graph windows per episode in graphs/, plus clip_maps/ mapping sampled graph-input frames to original video frames 7.5 MB
findingdory_psfr_cues.tar.gz FindingDory, full validation split; 100 episodes at 1 FPS Per-episode NPZ files in psfr_cues/, with scores of shape (T, 6) and hsv_hists of shape (T, 432) 91.6 MB
hourvideo_graphs.tar.gz HourVideo, development split; 50 videos 90 scene-graph windows per video in graphs/, including source-frame boundaries at 0.5 FPS 14.8 MB
hourvideo_psfr_cues.tar.gz HourVideo, development split; 50 videos at 0.5 FPS Per-video NPZ files in psfr_cues/, with scores of shape (T, 6) and hsv_hists of shape (T, 432) 59.0 MB
multihop_egoqa_hourvideo_disjoint_selector_training.tar.gz MultiHop-EgoQA, train, internal_dev, and filtered official_validation; all 50 HourVideo dev video UIDs excluded Canonical graph and QA/evidence tables in corpus/, exclusion list, corpus hashes, and train.sh for caption alignment and LLM-Selector training 330.3 MB

Each archive extracts into a directory named after the archive, containing a README, a validation report, and per-file checksums. Use the extracted graphs/ directories with --graph-dir or --graphs-dir, or place their contents at the paths shown in the examples below.

For FindingDory, a graph's frame_count counts sampled graph-input frames (up to 32 per clip). Its original_source_* fields and the PSFR cue rows refer to the full 1 FPS video; use the included clip maps to distinguish these index spaces. Both benchmarks' downstream PSFR selection uses cue columns [0,1,2,3,4]. The current PSFR CLI recomputes cues from video and does not automatically load the downloaded NPZ caches. These per-video caches also differ from the archive formats expected by the evolutionary-search evaluators.

The training archive physically removes HourVideo dev UIDs from every split and contains 310,770 graph windows and 11,477 QA rows. The released loader retains 10,387 selector-training examples; their order and the 20,000 caption-alignment examples match the original runtime exclusion procedure. Selector training uses train,internal_dev; caption alignment uses train only. Its official_validation split is also filtered and is not the full original MultiHop evaluation split. Follow the archive's README and train.sh: the filtered corpus has a new dataset hash, and the historical expected-removal counts of 12 videos and 160 training QA examples must not be applied again. Official MultiHop annotation JSON files, frozen embeddings, and model weights must be obtained separately. Recomputing embeddings after filtering changes batch composition and can change float16 values.

Official datasets and model weights

Input Official source and required split Access and usage
MultiHop-EgoQA Dataset repository, train_annotations.json and MultiHop-EgoQA.json; original video segments from Ego4D The dataset card labels its license ego4d. Follow Ego4D registration and video access terms. The released visual-feature archive is not a substitute for the RGB videos needed by graph extraction.
FindingDory Full dataset, complete validation split and full episode videos Apache 2.0 dataset card. Use the full validation release, not the 96-frame subsample. Expected validation annotation count: 5,870 questions.
HourVideo Version 1.0 dataset, v1.0_release/json/dev_v1.0_annotations.json, option images, and the 50 development video UIDs Apache 2.0 benchmark. Its repository is gated; accept its conditions and share contact information before downloading. Obtain source videos via Ego4D access. The benchmark forbids inclusion of its data in training corpora and requires its canary string to remain intact. Expected dev count: 1,182 MCQs.

Install huggingface_hub separately to use its hf download command. The files above can be fetched from their official dataset repositories after accepting any applicable terms. Verify the exact annotation files used for the paper with SHA-256:

hf download SurplusDeficit/MultiHop-EgoQA --repo-type dataset \
  --revision 4b73a71f3d0f74c1a78854d3d984c85b1024de1a \
  --include train_annotations.json --include MultiHop-EgoQA.json \
  --local-dir external_data/multihop
hf download yali30/findingdory --repo-type dataset \
  --revision 016160dd7ec1d0d23aa09bba7d48470e835030a0 \
  --include data/train-00000-of-00001.parquet \
  --include data/validation-00000-of-00001.parquet --include videos.zip \
  --local-dir external_data/findingdory/download
hf download HourVideo/HourVideo --repo-type dataset \
  --revision 48ac2d7b93eb1782154fad9c274cd0f1b4c78192 \
  --include 'v1.0_release/json/*' --include 'v1.0_release/navigation_images/*' \
  --include 'v1.0_release/spatial_layout_images/*' --include 'v1.0_release/video_uids/*' \
  --include 'prompts/baseline_evaluations/gemini-1.5-pro/qa_eval.yaml' \
  --local-dir external_data/hourvideo
sha256sum external_data/multihop/train_annotations.json \
  external_data/multihop/MultiHop-EgoQA.json \
  external_data/findingdory/download/data/validation-00000-of-00001.parquet \
  external_data/hourvideo/v1.0_release/json/dev_v1.0_annotations.json

Extract FindingDory's official videos.zip into external_data/findingdory/videos/ and check that the expected ep_<number>.mp4 files are present. For Ego4D RGB videos, register and obtain credentials through the official access process; the Ego4D CLI can fetch v2 metadata with ego4d --output_directory external_data/ego4d --datasets annotations --metadata --version v2, then fetch only required video UIDs with --datasets full_scale --video_uid_file <uid-file>. Resample or trim the obtained RGB videos into the benchmark source media listed below; verify their FPS with focusgraph video-manifest. MultiHop-EgoQA's feature archive is not RGB source media.

MultiHop-EgoQA/train_annotations.json       f899459abb53b44efd9a66e54e2264767fb0735c2b0ae2e7f017cb7425d7f89a
MultiHop-EgoQA/MultiHop-EgoQA.json           5066081d39ce80e53cf55dab9d5bb8804a111e725cb8c8fcdb00ffe44351602c
FindingDory/validation-00000-of-00001.parquet 5ba6ea653fecf33131b1f5f06b352356b38eff5a5780ab2dce02c8dba27a499e
HourVideo/dev_v1.0_annotations.json         e1af087df035d524ee64d92d34d2e81f29461fdeb5fa72cfda679d7c9909daf3

The pipeline uses answerdotai/ModernBERT-large at revision 45bb4654a4d5aaff24dd11d4781fa46d39bf8c13 for graph embeddings, Qwen/Qwen3.5-9B at revision c202236235762e1c871ad0ccb60c8ee5ba337b9a for the selector, and Qwen/Qwen2.5-VL-7B-Instruct for local QA. The HourVideo QA experiment used revision cc594898137f460bfe9f0759e9844b3ce807cfb5 of the QA model. Obtain these weights from their publishers under the applicable terms.

The trained FocusGraph selector checkpoints are not included, and download links are not yet available. FindingDory and HourVideo use separate checkpoints, as recorded in the protocol manifest.

Organizing the inputs

All commands accept explicit paths, so you can keep data wherever convenient. The examples use this layout:

external_data/
  findingdory/videos/ep_15.mp4
  findingdory/download/data/validation-00000-of-00001.parquet
  hourvideo/videos_0.5fps/<video_uid>.mp4
  hourvideo/graphs/<video_uid>.json
  hourvideo/v1.0_release/json/dev_v1.0_annotations.json
  multihop/train_annotations.json
  multihop/MultiHop-EgoQA.json
outputs/

HourVideo graph files can be downloaded above or generated during preprocessing. Each contains 90 windows, fps: 0.5, frame_count, and an original_source_start_frame and original_source_end_frame for every window. HourVideo assigns remainder frames to the earliest windows (video-manifest --window-policy balanced_front); FindingDory uses integer floor boundaries.

For single-video retrieval, provide a JSON array of 90 finite, nonnegative selector scores. An empty array ([]) means that the model produced no valid temporal-tag pair. Retrieval uses a strict threshold: a score equal to factor * max(score) is excluded.

Graph construction and frozen triplet features

The example below processes one FindingDory video that has already been sampled at 1 FPS. For HourVideo, use --expected-fps 0.5 --window-policy balanced_front when creating the manifest and --policy HV90 for C3. For 180-second MultiHop-EgoQA segments, use --expected-fps 5 and --policy W2.

The video manifest records the media SHA-256, FPS, all 90 clip boundaries, and source-frame indices. graph-inputs then creates the TSV used by the SG-Ego stages and a FindingDory clip map. Keep generated captions, graphs, and embeddings in an output directory outside the release.

focusgraph video-manifest --video external_data/findingdory/videos/ep_15.mp4 \
  --video-id ep_15 --split validation --expected-fps 1 \
  --output outputs/ep_15_manifest.json
focusgraph graph-inputs --manifest outputs/ep_15_manifest.json \
  --output-dir outputs/ep_15_inputs
focusgraph-c1 --manifest outputs/ep_15_inputs/segments.tsv \
  --sgego-root external_models/sg-ego --output-dir outputs/c1 --device cuda:0
focusgraph-c2 --manifest outputs/ep_15_inputs/segments.tsv \
  --sgego-root external_models/sg-ego --captions-dir outputs/c1 \
  --output-dir outputs/c2
focusgraph-c3 --manifest outputs/ep_15_inputs/segments.tsv \
  --sgego-root external_models/sg-ego \
  --input-dir external_data/findingdory/videos \
  --frame-graphs-dir outputs/c2/mapped --c2-success-dir outputs/c2/success \
  --output-dir outputs/c3 --policy FD90
focusgraph graph-embeddings --graphs outputs/c3/FD90/segments/ep_15.json \
  --model-revision 45bb4654a4d5aaff24dd11d4781fa46d39bf8c13 --dataset-hash <canonical-dataset-hash> --device cuda:0 \
  --output-dir outputs/ep_15_triplet_embeddings

graph-embeddings serializes each triplet as subject relation object, removes duplicates in first-occurrence order, and encodes the result with a 64-token limit. It saves L2-normalized float16 vectors, index and offset files, and metadata containing graph hashes and the ModernBERT revision.

For training, use --canonical-windows outputs/canonical/window_graphs.parquet as shown below. This preserves the source builder's row order. The default batch size is 512. Changing the order of texts or the batch size can change float16 results, so a cache generated from a subset of graphs need not match the corresponding rows of the full training cache. Metadata records the batch size, ordering, input hashes, model revision, and library versions. Use --index-only to check indexing without loading ModernBERT.

The SG-Ego stages record completed work with success markers, so interrupted runs can be resumed. Failures are written to each stage's output directory.

Training the GeLM selector

First, build the training corpus from the MultiHop scene graphs. Place the 5 FPS C3 W2 outputs under outputs/multihop_graphs/official_train/W2/segments/ and official_validation/W2/segments/. Each 180-second segment should have 90 two-second windows.

The segment manifest normalizes annotations and checks that training and validation videos do not overlap. The canonical corpus commands produce Parquet files and a dataset_hash.txt. Install the training dependencies with pip install -e '.[train]' in a CUDA PyTorch environment, and keep the generated corpus outside the release directory.

The commands below follow the HourVideo-disjoint recipe, using four A100 80 GB GPUs and ZeRO-3. This recipe excludes all 50 HourVideo development videos from both alignment and selector training. Before running it, create the exclusion file from the official UID list:

python -c 'from pathlib import Path; p=Path("external_data/hourvideo/v1.0_release/video_uids/dev.txt"); q=Path("external_data/hourvideo/dev_video_uids.txt"); q.write_text("\n".join(sorted(set(p.read_text().split())))+"\n")'
sha256sum external_data/hourvideo/dev_video_uids.txt

Then build the corpus, compute embeddings, and train the selector:

focusgraph-segment-manifest \
  --train-json external_data/multihop/train_annotations.json \
  --validation-json external_data/multihop/MultiHop-EgoQA.json \
  --ego4d-metadata external_data/ego4d/v2/ego4d.json \
  --output-dir outputs/multihop_manifest
focusgraph-canonical build \
  --manifest outputs/multihop_manifest/segments.parquet \
  --train-json external_data/multihop/train_annotations.json \
  --validation-json external_data/multihop/MultiHop-EgoQA.json \
  --graphs-root outputs/multihop_graphs --output-dir outputs/canonical
focusgraph-canonical freeze \
  --manifest outputs/multihop_manifest/segments.parquet \
  --train-json external_data/multihop/train_annotations.json \
  --validation-json external_data/multihop/MultiHop-EgoQA.json \
  --graphs-root outputs/multihop_graphs --output-dir outputs/canonical
focusgraph graph-embeddings \
  --canonical-windows outputs/canonical/window_graphs.parquet \
  --model-revision 45bb4654a4d5aaff24dd11d4781fa46d39bf8c13 \
  --dataset-hash <contents-of-outputs/canonical/dataset_hash.txt> \
  --device cuda:0 --output-dir outputs/triplet_cache
torchrun --standalone --nproc_per_node=4 \
  -m focusgraph.selector.train_focusgraph_caption_then_selector_dp \
  --caption-only --caption-examples 20000 --per-device-batch-size 2 \
  --accumulation 2 --seed 17 \
  --corpus-dir outputs/canonical --triplet-cache-dir outputs/triplet_cache \
  --exclude-video-uids-file external_data/hourvideo/dev_video_uids.txt \
  --expected-exclusion-list-size 50 --output-dir outputs/alignment
torchrun --standalone --nproc_per_node=4 \
  -m focusgraph.selector.train_focusgraph_gelm \
  --graph-input modernbert --caption-checkpoint outputs/alignment/caption_aligned \
  --corpus-dir outputs/canonical --triplet-cache-dir outputs/triplet_cache \
  --train-json external_data/multihop/train_annotations.json \
  --validation-json external_data/multihop/MultiHop-EgoQA.json \
  --exclude-video-uids-file external_data/hourvideo/dev_video_uids.txt \
  --expected-exclusion-list-size 50 --expected-excluded-train-video-uids 12 \
  --expected-excluded-train-examples 160 --stop-after-optimizer-steps 750 \
  --per-device-train-batch-size 1 --gradient-accumulation-steps 16 \
  --deepspeed "$(focusgraph config-path zero3_gelm)" --seed 17 \
  --output-dir outputs/gelm
focusgraph-evaluate-gelm --checkpoint outputs/gelm/final \
  --corpus-dir outputs/canonical --triplet-cache-dir outputs/triplet_cache \
  --train-json external_data/multihop/train_annotations.json \
  --validation-json external_data/multihop/MultiHop-EgoQA.json \
  --output-dir outputs/gelm_validation

The exclusion file must contain the 50 official HourVideo development UIDs, one per line. For the original experiment, its SHA-256 was 75c4d179bd50a1724a9dc4da94e5bfc31703115bc45599c816032032af27a292.

Alignment produces a projector state and metadata. GeLM training loads this state and saves resumable Hugging Face checkpoints and a final model. Its loss combines LM cross-entropy, temporal contrastive NCE, and weighted saliency BCE with a coefficient of 10. Evaluation saves teacher-forced losses, generated tagged answers, 90 clip scores, and temporal metrics.

Selecting clips and frames

Create a source-video manifest and check its 90 non-overlapping clip boundaries:

focusgraph video-manifest \
  --video external_data/findingdory/videos/ep_15.mp4 \
  --video-id ep_15 --split validation --expected-fps 1 \
  --output outputs/ep_15_manifest.json
focusgraph verify-splits --manifests outputs/*_manifest.json \
  --output outputs/split_check.json

Retrieve clip IDs from a selector score array:

focusgraph retrieve --scores external_data/scores_90.json --factor 0.3 \
  --output outputs/retrieval.json

PSFR computes motion cues over the full FindingDory source video at 1 FPS, then selects up to eight frames from the union of the retrieved clips:

focusgraph psfr --video external_data/findingdory/videos/ep_15.mp4 \
  --scores external_data/scores_90.json --factor 0.3 --budget 8 \
  --output outputs/findingdory_frames.json

For HourVideo, use the generated graph's exact source-frame boundaries. The source video must be sampled at 0.5 FPS. Empty retrieval falls back to the full video, as in the recorded experiment. For predictive questions, pass the first relevant timestamp as --predictive-cutoff-sec:

focusgraph hourvideo-psfr \
  --video external_data/hourvideo/videos_0.5fps/<video_uid>.mp4 \
  --graph external_data/hourvideo/graphs/<video_uid>.json \
  --scores external_data/scores_90.json --factor 0.3 --budget 8 \
  --output outputs/hourvideo_frames.json

The output records source-frame indices and timestamps. Both benchmarks use PSFR columns [0,1,2,3,4]: corner count, center corners, edge density, entropy, and low retention relative to the first frame. The protocol manifest records this configuration.

Batch selection and downstream QA

The inference commands below take C3 graphs and a trained GeLM checkpoint, then write the predictions used by frame selection. FindingDory records contain ordinal, ep_id, question, and similarity_scores_90; HourVideo records contain qid, video_uid, question, and similarity.scores_90.

Use the FindingDory checkpoint for the FindingDory results and the HourVideo-disjoint step-750 checkpoint for HourVideo. These weights must be supplied separately. The commands generate ModernBERT feature caches at the specified output paths; --modernbert-revision must match the weights used to create the training embeddings.

The HourVideo checkpoint directory must contain language_model/, focusgraph_modules.pt, and the tokenizer files. Its parent directory must also contain the original training run's focusgraph_gelm_config.json and trainer_state.json. The extractor verifies the exclusion-list checksum, absence of video overlap, graph input type, and optimizer step 750 from these files; copying the model weights alone is insufficient.

For the full FindingDory validation split, run selector inference, frame selection, and Qwen-2.5-VL-7B QA in sequence. This example uses retrieval factor 0.3 throughout, matching the corresponding paper row:

focusgraph-findingdory-infer \
  --checkpoint external_models/findingdory_selector/checkpoint_snapshot \
  --annotations external_data/findingdory/download/data/validation-00000-of-00001.parquet \
  --graph-dir external_data/findingdory/graphs \
  --embedding-cache outputs/findingdory_modernbert \
  --modernbert-revision 45bb4654a4d5aaff24dd11d4781fa46d39bf8c13 \
  --similarity-factor 0.3 --resume --output-dir outputs/findingdory_infer
focusgraph psfr-batch --benchmark findingdory \
  --predictions outputs/findingdory_infer/predictions.jsonl \
  --video-dir external_data/findingdory/videos \
  --factor 0.3 --budget 8 --output-dir outputs/findingdory_selection
focusgraph findingdory-qa \
  --annotations external_data/findingdory/download/data/validation-00000-of-00001.parquet \
  --selections outputs/findingdory_selection/selections.jsonl \
  --video-dir external_data/findingdory/videos \
  --model-revision cc594898137f460bfe9f0759e9844b3ce807cfb5 \
  --output-dir outputs/findingdory_qa

The protocol manifest records the release settings and reported paper metrics. This example covers the Qwen-2.5-VL-7B row; other published rows may use different settings.

For the HourVideo development split, use the corresponding commands:

focusgraph-hourvideo-infer \
  --checkpoint external_models/hourvideo_disjoint_selector/final \
  --annotations external_data/hourvideo/v1.0_release/json/dev_v1.0_annotations.json \
  --graph-dir external_data/hourvideo/graphs \
  --denylist external_data/hourvideo/dev_video_uids.txt \
  --modernbert-revision 45bb4654a4d5aaff24dd11d4781fa46d39bf8c13 \
  --output-dir outputs/hourvideo_infer
focusgraph psfr-batch --benchmark hourvideo \
  --predictions outputs/hourvideo_infer/predictions.jsonl \
  --video-dir external_data/hourvideo/videos_0.5fps \
  --graphs-dir external_data/hourvideo/graphs \
  --annotations external_data/hourvideo/v1.0_release/json/dev_v1.0_annotations.json \
  --factor 0.3 --budget 8 --output-dir outputs/hourvideo_selection
focusgraph hourvideo-qa \
  --annotations external_data/hourvideo/v1.0_release/json/dev_v1.0_annotations.json \
  --selections outputs/hourvideo_selection/selections.jsonl \
  --video-dir external_data/hourvideo/videos_0.5fps \
  --option-images-dir external_data/hourvideo/v1.0_release \
  --prompt-file external_data/hourvideo/prompts/baseline_evaluations/gemini-1.5-pro/qa_eval.yaml \
  --output-dir outputs/hourvideo_qa

Add --dry-run to either QA command to validate IDs, media paths, question counts, required prompts, and option images without loading the model.

QA results are saved atomically after each question, and rerunning the command resumes from completed predictions. HourVideo applies the predictive cutoff before QA and letterboxes source frames and option images to 512×384.

FindingDory inference reports temporal retrieval metrics against the full validation frame references; its QA command reports relaxed accuracy. HourVideo inference records coverage and scores, and its QA command reports multiple-choice accuracy. The HourVideo result in the paper averages seeds 2026, 2027, and 2028.

Evolving the PSFR selection program

The search code is adapted from the PSFR examples in the original OpenEvolve project. The release preserves their evaluators, subprocess runners, initial programs, prompts, and search settings. Paths are supplied explicitly, and API credentials are read from the environment. The exact source files and changes are recorded in the provenance manifest.

Three search variants are included:

Variant Frame budget Candidates Objective in the available source
findingdory 96 Per-question candidate indices supplied in model_outputs.json Intersection with a runtime penalty
gens 96 All frames of each video Weighted intersection with a runtime penalty
combined 16 All frames of each video, evaluated on FindingDory and GenS Normalized combination of intersection, inclusion, and F-score, then a geometric mean across datasets

All three use a 15-second subprocess timeout, six input score channels, and 432-dimensional HSV histograms. The current standalone runners hard-code 96 frames; the combined runner uses 16. These are the interfaces of the original search programs. The downstream QA selector has a separate five-channel adapter, described above.

Install the optional search dependencies:

pip install -e '.[evolution]'
focusgraph-psfr-evolution --help
focusgraph-psfr-evolution paths --variant findingdory

For FindingDory, provide the original-format feature archive and the two aligned JSON files. See Search inputs and provenance for their schemas and the inputs needed by the other variants. No features or annotations are distributed with the code.

First evaluate the original uniform initial program locally. This command does not contact an LLM:

focusgraph-psfr-evolution evaluate --variant findingdory \
  --fd-scores external_data/evolution/findingdory/scores_old.npz \
  --fd-segments external_data/evolution/findingdory/model_outputs.json \
  --fd-gt external_data/evolution/findingdory/gt_answers.json \
  --output outputs/evolution/uniform.json

To start a search, set DEEPSEEK_API_KEY in your environment and run:

focusgraph-psfr-evolution search --variant findingdory \
  --fd-scores external_data/evolution/findingdory/scores_old.npz \
  --fd-segments external_data/evolution/findingdory/model_outputs.json \
  --fd-gt external_data/evolution/findingdory/gt_answers.json \
  --output-dir outputs/evolution/findingdory

Add --check-only to check the input paths and configuration without making API requests. Repeating a search command resumes its latest checkpoint. Use --program, --config, or --checkpoint to supply an alternative initial program, configuration, or checkpoint. The output directory also records hashes of the supplied inputs and program.

The findingdory_inclusion_785.py program is included in src/focusgraph/psfr/search/programs/. Running the downstream PSFR selector does not require OpenEvolve or an LLM service.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages