2015–2016
Permanent URI for this collection
Browse
Browsing 2015–2016 by Author "Bharath, B R"
Now showing 1 - 1 of 1
Results Per Page
Sort Options
Item Automation Project ( Machine Learning ) Data Extraction From PDF Documents(2016-12-01T13:27:57Z) Bharath, B RAutomation Project (Machine learning)-Data Extraction from PDF Documents Aim of the Project The project is amid to develop a system to extract required financial data from the PDF or HTML files. These files consist of financial data of current year. Data are provided by companies as part of their financial statement which is not available in the database Technical Details This project deals with several different markets that belongs to many country’s Each country market as a rule File that specifies criteria and filter to be applied. The rules are stored in a xls file depending on the country rules are applied It removes unwanted columns from the input xml file. The types of financial statements are stored in SectionTag file which consist of statement type and statement Code used to Find StatementInstance. The Data extracted from the HTML Files are refined and required footnote Description is fetched to a footnote Description List and using FincancialConceptCode and FincancialConceptDescription in Database it is matched with footnote_Description that is in the HTML table Collection to find the required data.