Project Collection Workspace (reanimator.collection)#
The ProjectCollection class provides a memory-efficient workspace manager that streams documents directly to disk without keeping large collections in RAM.
ProjectCollection#
class ProjectCollection(
project_dir: Union[str, Path],
arxiv_ids: Optional[List[str]] = None,
dois: Optional[List[str]] = None,
local_pdfs: Optional[List[Union[str, Path]]] = None,
topics: Optional[Union[List[str], List[Dict[str, Any]]]] = None,
name: str = "reanimator_project",
)
Workspace Layout#
project_dir/
├── manifest.json # Workspace manifest tracking inputs & status
├── pdfs/ # Raw PDF files
├── parsed_docs/ # Parsed Document JSON artifacts
├── figures/ # Extracted figure images
├── cache/ # Disk-backed VLM response cache
└── relevance/ # Synthetic judgements & TREC qrels files
Methods#
process(pipeline: ReanimatorVLM, email: Optional[str] = None, verbose: bool = True) -> List[str]#
Batch process all collection inputs via VLM pipeline. Saves each parsed Document directly to parsed_docs/{doc_id}.json on disk and releases RAM immediately.
iter_documents() -> Iterator[Document]#
Generator yielding Document instances lazily from disk.
get_document(doc_id: str) -> Optional[Document]#
Load a single Document directly from disk by ID.
run_relevance_assessment(evaluator: RelevanceEvaluator, modality: str = "tables", verbose: bool = True) -> List[Judgement]#
Run synthetic relevance assessment across collection documents lazily from disk. Saves judgements_{modality}.json and exports TREC qrels_{modality}.txt.