Description
Role Overview
As a Software Architect, you will own the design and implementation of the high-level software layers that connect open-source big data engines such as Apache Spark to DualBird's custom hardware accelerators. This includes the Spark plugin, query planning and optimization, execution orchestration, and the interfaces between the engine, the cloud environment, and the accelerated runtime.
This is a hands-on architecture role. You will set technical direction and make system-wide design decisions, but you will also write a significant share of the code yourself while driving complex features end-to-end. You will work closely with the compiler, runtime, hardware, and infrastructure teams to ensure the full stack delivers predictable, out-of-the-box performance on large-scale customer workloads running on AWS.
Responsibilities
· Design and evolve the architecture of DualBird's Spark integration and query processing layers.
· Lead the design and implementation of new features across planning, optimization, and execution, from concept to production.
· Extend support for Spark and SQL semantics, expressions, and operators.
· Own correctness, compatibility, and performance across Spark versions and AWS data platforms (EMR, EKS, Glue, S3).
· Review designs and code, mentor engineers, and shape engineering practices across the team.
· Work with customers' real workloads to identify gaps, prioritize, and deliver improvements.
Requirements
Required Qualifications
· Bachelor's degree in Computer Science or a related field from a leading university; advanced degree an advantage.
· 8+ years of software development experience, including several years in a technical lead or architect capacity.
· Proven track record of designing large, complex software systems and delivering them hands-on in production.
Preferred Qualifications
· Experience with query engines, database internals, or query optimization.
· Strong proficiency in Scala and/or Java, and comfort working in a mixed JVM / native (C++) codebase.
· Familiarity with columnar and vectorized execution (e.g., Velox, Arrow, Gluten, Photon) and columnar storage formats such as Parquet.
· Experience with AWS data services (EMR, EKS, S3, Glue) and cloud-native deployment.
· Background in performance analysis and benchmarking of large-scale data pipelines.
· Experience with hardware-accelerated or heterogeneous computing systems.
· Comfortable working across the stack with runtime and hardware teams.
Send Application
Submit the form below to send your application.
Transform your data infrastructure performance with a few clicks
Zero risk, zero effort, incredible results.

