About this role
Work Flexibility: Hybrid
What You will do:
• Analyze incoming order documents and map out their formats, fields, and edge cases. • Evaluate OCR and document-extraction tools (e.g. Tesseract, PaddleOCR, layout-aware models like LayoutLM/LiLT) and help choose the right approach. • Build a pipeline that extracts structured order data (customer, part numbers, quantities, prices, dates, PO references) from documents. • Implement validation against master data and a review step that flags uncertain results for a human to check. • Integrate the output with existing data infrastructure for downstream reporting and order entry. • Document your work and support handover to the team.
What you will need:
Required Qualification:
• Working knowledge of Python (data processing, scripting). • Masters degree in Engineering passing out in 2025 or 2026 • Interest in (or some exposure to) machine learning, computer vision, or NLP. • Basic SQL and comfort with structured data. • An analytical mindset and strong attention to data quality • Basic knowledge of UI-based automations (eg. UiPath) and/or API-based automations (eg. PowerAutomate)
Preferred Qualification:
• Experience with OCR libraries or document-AI models. • Familiarity with cloud data platforms and pipeline tooling. • Version control (Git) and collaborative development experience.
Travel Percentage: 0%