Skip to content

Feature/table download - #130

Merged
Jordan-Leis merged 14 commits into
mainfrom
feature/table-download
Jun 19, 2026
Merged

Jordan-Leis merged 14 commits into
mainfrom
feature/table-download

Conversation

@Jordan-Leis

Copy link
Copy Markdown
Collaborator

Added the ability for users to download tables and extract them for use in excel and other apps

Sam120204 and others added 8 commits May 14, 2026 23:05
… processing, and LLM layer documentation

Introduced comprehensive documentation covering the backend architecture, API surface, data models, Pydantic schemas, document processing workflows, and the LLM provider layer. This includes detailed descriptions of application entry points, runtime layers, package responsibilities, and main data flows, enhancing the overall understanding of the backend system.
… data models, document processing, LLM routing, extraction, evaluation, session sharing, and authentication. Added diagrams to improve understanding of system interactions and data flows.
…xisting documentation with additional links for better navigation. Added a new class reference appendix for detailed field-level insights on ORM models, Pydantic schemas, and service classes.
Introduced a comprehensive README file detailing the Science-GPT web application, its purpose, target users, functionality, testing methodology, and evaluation results. This documentation aims to provide a clear overview of the tool's capabilities and the benefits it offers to scientific reviewers in processing large documents.
- Introduced `14-batch-results.md` detailing the Batch Results page, including UI sections, table structure, fuzzy search, row detail modal, Excel export, state management, and API calls.
- Updated `README.md` to include an overview of the Batch Results feature and its integration within the Science-GPT frontend.
- Created `component-index.md` to catalog shared components used across multiple pages.
- Added `hooks-contexts.md` to document custom hooks and context providers utilized in the frontend.
- Defined key TypeScript interfaces in `types-interfaces.md` for better understanding of the data structures used.
- Established a glossary to clarify project-specific terms and Azure services relevant to the Science-GPT project.
- TablesGallery: download individual tables as HTML or all tables
  bundled into a single file for Excel import
- auth-service: conditional SSL for local vs Azure Postgres; run
  Better Auth DB migrations on startup
Comment thread auth-service/src/index.ts Fixed
Replace .includes("azure.com") substring match with URL parsing so
the azure.com check only matches the hostname (exact or subdomain),
and sslmode is checked via searchParams rather than a full-string
include. Resolves CodeQL js/incomplete-url-substring-sanitization.

@Sam120204 Sam120204 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OECD GD 150 (2018) Revised Guidance Document 150 on Standardised Test Guidelines for Evaluating Chemicals for Endocrine Disruption-CURRENT.pdf I tested with this large file, it contains 266 tables, and the issue right now is that the backend is busy downloading the tables so everything is blocking and showing timeout (like displaying pdf in the frontend). im thinking about maybe implmenting batching downloading tables. but i tried other smaller files, and download is working. so i guess optimizing the download for larger files would be the priority. But overall looks good :)

@Sam120204

Copy link
Copy Markdown
Contributor

also remember to merge from the origin/main to this branch :)

@spencer-crook
spencer-crook self-requested a review June 11, 2026 19:26

@spencer-crook spencer-crook left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me!

Jordan-Leis and others added 4 commits June 12, 2026 18:23
Documents with 266+ tables caused the PDF viewer to block on load and
"Download All" to time out. Three targeted fixes:

- Add GET /documents/{id}/tables/download endpoint that fetches all table
  HTML files concurrently via asyncio.gather() and returns a single ZIP,
  collapsing 266 HTTP round-trips into one (32x faster in benchmarks)
- Wrap tmp_file.read_bytes/write_bytes in asyncio.to_thread() so /tmp
  cache I/O no longer blocks the async event loop during burst requests
- Lazy-load TablePreview cards with IntersectionObserver so only visible
  cards fetch their HTML on render (eliminates 266 simultaneous requests
  that were starving the PDF viewer of browser connections)
@Jordan-Leis
Jordan-Leis requested a review from Sam120204 June 12, 2026 18:40
@Jordan-Leis
Jordan-Leis merged commit e99ee9b into main Jun 19, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants