Apply now »

Title:  Director, Software Engineer Data Processing Engine Orchestration

Location: 

San Jose, CA, US, 95128

Requisition ID:  136550

Job Summary

DataPelago is at the forefront of revolutionizing data processing for traditional analytics and cutting-edge GenAI preprocessing. We are building an innovative data processing engine that is transforming how Apache Spark, Apache Flink, Ray and others operate on diverse, large-scale data. Our team of engineers drive and adopt advances in hardware-accelerated computing, parallel processing of large-scale data, query optimization, distributed systems, compilers, machine learning, and cloud-native computing. We are looking for specialists to join our engineering team and shape the future of accelerated data processing.

 

DataPelago Nucleus is a universal data processing engine that is designed to accelerate the processing of diverse data – structured through unstructured – with any parallel processing framework – e.g., Apache Spark – on any infrastructure – including vectorized CPU and GPU. DataPelago Accelerator for Spark (DPA-S), our product based on DataPelago Nucleus, is running large-scale production applications of many globally renowned customers.

 

The Opportunity:

 

As senior development leader of DataPelago data processing engine orchestration, you will spearhead the development and delivery of innovative capabilities in the planning, optimization, and coordination of complex workloads on large-scale distributed systems. The orchestration layer drives distributed execution plans adapted to the capabilities of the processing elements in the infrastructure, towards minimizing execution duration and cost. You will be responsible for growing and advancing the development team, development methodologies, and product capabilities.

a key individual contributor in adopting and advancing capabilities of open-source software (OSS) platforms such as Apache Gluten, Velox, Apache Spark, and Apache Flink in the context of DataPelago’s data processing engine. You will enhance functional breadth, performance, scale, and reliability of the DataPelago engine through downstream and upstream contributions. You will have the opportunity to engage with community working on these platforms. This is a unique opportunity to make a significant impact on a category-defining product and work with a talented team of engineers.

What You'll Do:

 

  • Lead and grow a talented team of engineers with expertise in query planning, query optimization, plan execution orchestration for massively parallel data processing engines.
  • Drive development with technical leads to implement and deliver industry-leading capabilities in data processing engine orchestration.
  • Establish metrics including for product quality, performance, scale, and reliability and lead the team in ensuring these metrics are continuously improved and achieved in customer deployments.
  • Continuously improve software development lifecycle practices leveraging the latest AI advances for higher code quality, release velocity, and productivity.
  • Execute efficiently to fulfill product roadmap aligned with organization goals and priorities.
  • Lead by example and foster a high-performance, high-integrity, collaborative work culture.

What You'll Bring:

 

  • Bachelor’s degree in computer science or a related field. Master’s or PhD preferred.
  • 12+ years of production software development experience with 7+ years of experience managing software development teams.
  • 10+ years of experience in planning, developing, launching, and supporting core components of massively parallel data processing platforms (e.g., query engines, cloud data warehouses, lakehouse engines) for production deployment.
  • Technical and process rigor in meeting and exceeding data accuracy, high performance, reliability, and security requirements of enterprise data processing products.
  • Ability to attract, retain, and grow exceptional talent in data processing engine orchestration.
  • Deep technical experience in one or more of database systems, data warehouses, query optimization, query engines. Experience with one or more of Apache Spark, Apache Flink, Apache Gluten, Velox, Apache DataFusion, Spark RAPIDS preferred.
  • Solid understanding of the architecture, deployment, and operations of data processing platforms in applications such as data engineering, data preparation, and analytics.
  • Excellent judgment on technology, software development, AI adoption, and people management. Strong leadership skills in all these areas.
  • Good skills engaging with customers, product management, sales, and executive management.

Compensation:
The target salary range for this position is 227,800 - 338,800 USD. The salary offered will be determined by the candidate's location, qualifications, experience, and education and may be outside of this range. Final compensation packages are competitive and in line with industry standards, reflecting a variety of factors, and include a comprehensive benefits package. This may cover Health Insurance, Life Insurance, Retirement or Pension Plans, Paid Time Off, various Leave options, Performance-Based Incentives, employee stock purchase plan, and/or restricted stocks (RSU’s), with all offerings subject to regional variations and governed by local laws, regulations, and company policies. Benefits may vary by country and region, and further details will be provided as part of the recruitment process. 


Nearest Major Market: San Jose
Nearest Secondary Market: Palo Alto

Job Segment: Cloud, Software Engineer, Process Engineer, Computer Science, Outside Sales, Technology, Engineering, Sales

Apply now »