Home /
[Remote] Big Data Dev/Spark Scala Engineer
– Amazon Store
[Remote] Big Data Dev/Spark Scala Engineer
Other Jobs To Apply
No other job posts for this day.
Note: The job is a remote job and is open to candidates in USA. Tata Consultancy Services is seeking a Big Data Dev/Spark Scala Engineer to design, develop, and maintain large scale Spark applications. The role involves building and operating streaming heavy data pipelines, implementing stateful streaming patterns, and ensuring data quality and consistency.
Responsibilities
Design, develop, and maintain large scale Spark applications using Scala and PySpark
Build and operate streaming heavy data pipelines using Kafka and Spark Structured Streaming
Implement stateful streaming patterns including windowing, watermarking, late data handling, and checkpointing
Develop robust event replay and reprocessing workflows using Kafka offsets and partitions
Build ingestion and routing flows using Apache NiFi, including Kafka based ingestion patterns
Implement end to end ETL/ELT pipelines with strong emphasis on low latency, fault tolerance, and scalability
Optimize Spark jobs through partitioning strategies, memory tuning, shuffle optimization, and efficient data formats
Integrate Spark worklo with distributed object storage systems such as Apache Ozone and Ceph
Ensure data quality, consistency, and auditability through validation, reconciliation, and metadata capture
Collaborate with platform, infrastructure, and operations teams on production readiness and capacity planning
Support production systems, including monitoring, incident analysis, and root cause resolution
Contribute to reusable frameworks, coding standards, and engineering best practices
Participate in architecture reviews, code reviews, and technical documentation
Skills
Experience Required - 7+ Years
Experience with Apache Ozone and/or Ceph as storage backends for analytics worklo
Experience implementing exactly once / at least once streaming semantics
Strong background in Spark performance tuning (CPU, memory, I/O, shuffle)
Experience supporting mission critical production systems with strict SLAs
Familiarity with CI/CD pipelines and automated testing for data applications
Experience designing observability for streaming systems (lag, throughput, backpressure)
Languages: Scala, Python (PySpark), SQL
Big Data: Apache Spark (Core, SQL, Structured Streaming)
Tata Consultancy Services is a business solutions company that specializes on information technology services and consulting. It is a sub-organization of Tata Group. It was founded in 1968, and is headquartered in Mumbai, Maharashtra, IND, with a workforce of 10001+ employees. Its website is
Company H1B Sponsorship
Tata Consultancy Services has a track record of offering H1B sponsorships, with 1844 in 2026, 7880 in 2025, 9690 in 2024, 8537 in 2023, 11159 in 2022, 9813 in 2021, 11984 in 2020. Please note that this does not guarantee sponsorship for this specific role.