Associate Data Engineer

Associate Data Engineer / Data Engineer

Experience: 2–4 Years

About the Team

Our Data Engineering team is the backbone of PayU’s data-driven organization. We design, build, and operate scalable data platforms that support critical business decisions across multiple domains.

Working with large and diverse data volumes, the team enables reliable data movement from operational systems to data warehouses, data lakes, lakehouses, and analytics platforms. We collaborate closely with business and technology stakeholders to build fit-for-purpose data solutions that address evolving organizational needs.

About the Role

As an Associate Data Engineer / Data Engineer, you will contribute to the design, development, and maintenance of scalable data pipelines and processing solutions. You will work with batch and real-time data platforms, support data warehouse and lakehouse initiatives, and help ensure the performance, quality, reliability, and accessibility of enterprise data.

This role provides an opportunity to work with modern cloud and open-source technologies while developing expertise across data warehousing, data lakes, lakehouses, and analytical serving platforms.

Key Responsibilities

  • Work with data engineers, architects, product teams, and business stakeholders to understand data requirements and deliver reliable technical solutions.
  • Participate in requirements gathering, solution design, development, testing, deployment, and support for data engineering initiatives.
  • Design, develop, and maintain reliable near-real-time (NRT) data pipelines using technologies such as Apache Kafka, Apache Flink, Python, and SQL.
  • Build and support NRT data processing workflows that enable timely analytics, monitoring, reporting, and downstream business use cases.
  • Develop and maintain Flink-based data lake solutions for processing high-volume, continuously changing data.
  • Work with Apache Iceberg as the table and data format for building scalable, reliable, and queryable data lakehouse platforms.
  • Contribute to data ingestion, transformation, schema evolution, partitioning, compaction, and data lifecycle management within Iceberg-based data lakes.
  • Develop and maintain analytical datasets, data models, and DataMart solutions for reporting and business intelligence.
  • Build and optimize data processing and query workloads using open-source technologies such as Apache Spark, Trino, and StarRocks.
  • Support low-latency analytical serving and interactive query use cases using open-source query engines such as StarRocks and Trino.
  • Use Apache Spark for distributed data processing, transformations, data preparation, and large-scale analytical workloads.
  • Write efficient SQL queries and contribute to query optimization, data partitioning, and performance improvements across data platforms.
  • Implement data validation, quality checks, monitoring, alerting, and observability across NRT pipelines and data lake workloads.
  • Troubleshoot data pipeline and production issues, perform initial root-cause analysis, and support timely incident resolution.
  • Write maintainable, testable, and well-documented code.
  • Participate in code reviews and follow engineering standards and best practices.
  • Contribute to technical documentation, operational runbooks, and continuous improvement initiatives.
  • Work effectively in an Agile and collaborative engineering environment.

Required Qualifications and Experience

  • 2–4 years of professional experience in Data Engineering, Software Engineering, Analytics Engineering, or a related discipline.
  • Strong programming and analytical skills, with hands-on experience in Python and SQL.
  • Good understanding of data structures, algorithms, relational databases, and software development principles.
  • Understanding of distributed systems, data processing, and large-scale data architectures.
  • Experience building, maintaining, or supporting ETL and ELT data pipelines.
  • Experience working with data warehouses, data lakes, or lakehouse platforms.
  • Understanding of data modeling, dimensional modeling, and DataMart development.
  • Experience working with batch processing; exposure to streaming or event-driven data pipelines is an advantage.
  • Understanding of SQL query optimization, data partitioning, indexing, and performance tuning.
  • Experience working with relational databases and familiarity with NoSQL databases.
  • Hands-on experience with cloud platforms, preferably AWS, including one or more of the following:
    • Amazon S3
    • Amazon EMR
    • Amazon Redshift
    • Amazon MSK
  • Practical experience with some of the following technologies:
    • Apache Spark / PySpark
    • Apache Airflow
    • Apache Kafka
    • Debezium
    • Apache Flink
    • Apache Iceberg
    • Trino
    • StarRocks
    • ClickHouse
    • Hadoop
  • Experience with Git, code reviews, testing, and CI/CD practices.
  • Kafka/kafka connect, debezium/maxwell. 
  • Binlog/wal/gtid log replication
  • Ability to troubleshoot issues, analyze problems, and communicate solutions clearly.
  • Strong communication and collaboration skills.

Good to Have

  • Experience with Change Data Capture, database replication, or incremental data processing.
  • Familiarity with data quality frameworks, pipeline monitoring, and observability practices.
  • Exposure to containerization technologies such as Docker and Kubernetes.
  • Familiarity with infrastructure automation and DevOps practices.
  • Experience working with on-premises or hybrid data platforms.
  • Exposure to Apache Iceberg, lakehouse architecture, or open table formats.
  • Experience working in fintech, payments, banking, or other high-volume transaction environments.
  • Familiarity with NoSQL databases such as Cassandra, MongoDB, or Redis.
  • Understanding of Agile development methodologies and engineering best practices.

What We Offer

  • A positive, collaborative, and get-things-done workplace.
  • A dynamic and constantly evolving environment where adaptability is valued.
  • An inclusive culture that encourages diverse perspectives and open communication.
  • Opportunities to work on modern data technologies at global scale.
  • The opportunity to learn and innovate in an agile fintech environment.
  • Access to 5,000+ training courses available anytime and anywhere through leading learning partners such as Harvard, Coursera, and Udacity.
  • Opportunities to work closely with experienced engineers and grow across data engineering, cloud, and distributed systems.

About PayU

At PayU, we are a global fintech investor with a vision to build a world without financial borders, where everyone can prosper. We provide people in high-growth markets with the financial services and products they need to thrive.

Our expertise across 18+ high-growth markets enables us to expand access to financial services. This drives everything we do—from investing in technology entrepreneurs and offering credit to underserved individuals, to helping merchants buy, sell, and operate online.

As part of Prosus, one of the world’s largest technology investors, PayU combines global reach, deep expertise, and the ability to create meaningful impact. Learn more at www.payu.com.

Our Commitment to Diversity and Inclusion

PayU is committed to building a diverse, inclusive, and safe workplace where every individual feels respected, valued, heard, and empowered to succeed.

As a global, multicultural organization, we welcome people from diverse backgrounds, identities, experiences, and perspectives. Our leaders are committed to fostering a transparent, flexible, and equitable work culture where everyone has the opportunity to grow and contribute.

PayU has zero tolerance for discrimination or prejudice based on race, ethnicity, gender, color, religion, disability, sexual orientation, gender identity, or any other personal characteristic.