Project Challenges

RPA is a Python-powered, end-to-end, automated, web-based data pipeline project capable of gathering and mining data from web-based scrapers, cleaning it according to an accepted structure, uploading it to the database, and then applying business logic to make necessary changes.

Project Solution

Real-time Dashboard and Reporting

Built web data scrapers for all types of websites (300+)

Data Enrichment and Transformation according to the data structure and requirements

Data uploading system for validation, comparison, and insertion

Web-based Data Pipeline Flow monitoring and status updates with error handling

Multiple instances (micro-services) for process management and control

Email integration for process compilation updates

Tor and VPN-based scrapers for advanced scraping control (speed and blocking issues)

Process job scheduling, file operation servers, logs, and reports, among other features

User Access Control and rule-based Security Measures

Technical Information

  • Web Server :  Windows Server IIS and Docker
  • Database :  MS SQL Server 2014+ , MYSQL
  • Programming languagess :   .NET core 5 ,C#, Python 3.10
  • Client Side Technologies :  React Js , HTML, CSS3, SCSS, Bootstrap, Java Script, flask 2.2.3
  • Version Control :  Git , SVN
  • Architecture :  Distributed and repository pattern
  • Other Tools :  TOR, VPN, Container