Title: Sr. Software Engineer, Data Processing OSS
San Jose, CA, US, 95128 US
Job Summary
DataPelago is at the forefront of revolutionizing data processing for traditional analytics and cutting-edge GenAI preprocessing. We are building an innovative data processing engine that is transforming how Apache Spark, Apache Flink, Ray and others operate on diverse, large-scale data. Our team of engineers drive and adopt advances in hardware-accelerated computing, parallel processing of large-scale data, query optimization, distributed systems, compilers, machine learning, and cloud-native computing. We are looking for specialists to join our engineering team and shape the future of accelerated data processing.
DataPelago Nucleus is a universal data processing engine that is designed to accelerate the processing of diverse data – structured through unstructured – with any parallel processing framework – e.g., Apache Spark – on any infrastructure – including vectorized CPU and GPU. DataPelago Accelerator for Spark (DPA-S), our product based on DataPelago Nucleus, is running large-scale production applications of many globally renowned customers.
The Opportunity:
As a Senior/Principal Data Processing Engineer - OSS, you will be a key individual contributor in adopting and advancing capabilities of open-source software (OSS) platforms such as Apache Gluten, Velox, Apache Spark, and Apache Flink in the context of DataPelago’s data processing engine. You will enhance functional breadth, performance, scale, and reliability of the DataPelago engine through downstream and upstream contributions. You will have the opportunity to engage with community working on these platforms. This is a unique opportunity to make a significant impact on a category-defining product and work with a talented team of engineers.
What You'll Do:
- Architect: Influence the architecture of how our data processing engine interfaces with open-source platforms and engines.
- Design: Lead design of functional and performance enhancements to open-source platforms such as Apache Gluten and Velox and their integration with our data processing engine.
- Core Development: Individually design, implement, test, optimize, and maintain components of the data processing engine.
- Innovation and Differentiation: Analyze technology roadmap of Apache Gluten, Velox, and equivalent platforms and identify opportunities for our engine to enhance technology and product leadership.
- Collaboration: Partner effectively with engineering, product management, open-source community and customer success teams.
- Continuous Improvement: Foster best practices in design and code reviews, testing, CI/CD, and issue resolution to maintain highest product quality, security, efficiency, & productivity.
What You'll Bring:
- 12+ years of relevant experience Bachelor's degree in Computer Science, or a related field OR a Master's degree in Computer Science.
- 3+ years of deep technical experience in developing core components of Apache Spark, Apache Flink, Apache Doris, Apache Gluten, Velox, Apache DataFusion, Apache DataFusion Comet, or equivalent platforms designed for large-scale data processing.
- 3+ years of deep technical experience in instrumenting, analyzing, and optimizing the performance of data processing engine components on benchmark and customer workloads.
- Sound knowledge of the architecture and internal operation of one or more of Apache Spark, Apache Flink, Presto/Trino.
- Demonstrated experience in the design, development, and successful release of high-performance data processing engines for large production deployments.
- Exceptional programming skills in C, C++, and Java.
- Extensive development experience in Linux environments.
- Strong analytical and problem-solving skills with a passion for performance optimization.
Compensation:
The target salary range for this position is 196,350 - 292,600 USD. The salary offered will be determined by the candidate's location, qualifications, experience, and education and may be outside of this range. Final compensation packages are competitive and in line with industry standards, reflecting a variety of factors, and include a comprehensive benefits package. This may cover Health Insurance, Life Insurance, Retirement or Pension Plans, Paid Time Off, various Leave options, Performance-Based Incentives, employee stock purchase plan, and/or restricted stocks (RSU’s), with all offerings subject to regional variations and governed by local laws, regulations, and company policies. Benefits may vary by country and region, and further details will be provided as part of the recruitment process.
Nearest Major Market: San Jose
Nearest Secondary Market: Palo Alto
Job Segment:
Open Source, Cloud, Test Engineer, Testing, Senior Product Manager, Technology, Engineering, Operations