I need a streamlined way to turn a large collection of PDFs into a searchable database focused on data extraction and analysis. The core requirement is to pull every piece of text data from each file—headings, body copy, footnotes, everything—and store it in a structured repository that I can query later.
Here’s what I’m looking for:
1. Ingestion & Parsing • A script or small utility that ingests bulk PDFs from a folder or S3 bucket. • Automatic OCR fallback for any scanned pages (Tesseract or a comparable engine). • Reliable parsing of the extracted text using Python, Apache Tika, pdfminer-six, or a tool you recommend.
2. Storage Design • A relational database (PostgreSQL/MySQL) or a text-search solution such as Elasticsearch—whichever best supports fast full-text queries. • Each document should be tagged with filename, date added, page numbers, and any basic metadata found in the PDF header.
3. Query & Export • Simple CLI or lightweight web interface where I can search keywords, filter by metadata, and export results to CSV for further analysis. • Basic analytics endpoints (word frequency, document counts) would be a plus.
4. Documentation & Handoff • Clear setup instructions, dependency list, and commented code. • A short README that shows example import, search, and CSV export commands.
I’m comfortable running this on a Linux environment and can provide sample PDFs immediately. The budget targets a functional prototype with clean, maintainable code rather than a polished enterprise UI, so focus on core extraction accuracy and searchable storage. If you’ve built similar ETL pipelines or OCR workflows, let me know—I value proven experience and concise solutions.
E-commerce Development for Online Portrait Art Store Category: Content Management System (CMS), Digital Marketing, ECommerce, HTML, SEO, Shopping Cart Integration, Web Development, Web Design Budget: $250 - $750 USD
02-Nov-2025 22:51 GMT
Large Shaded Beach Cabana Design Category: 3D Design, 3D Modelling, 3D Rendering, Architecture, AutoCAD, Building Design, CAD / CAM, Design, Drafting, SketchUp Budget: $30 - $250 AUD