GCP Data Architect interview Q&A – Part 2: pipelines, streaming & governance
1. Dataflow, Dataproc, Data Fusion or Composer – how do you choose?
- Dataflow – serverless Apache Beam for batch and streaming with autoscaling; default for new pipelines.
- Dataproc – managed Spark/Hadoop; best for existing Spark jobs and lift-and-shift (also Serverless for Spark).
- Data Fusion – visual, low-code pipelines.
- Cloud Composer – managed Apache Airflow for orchestration, not heavy processing.
- Dataform – SQL-based transformations inside BigQuery.
2. What delivery guarantees does Pub/Sub give?
At-least-once delivery by default, so consumers must be idempotent. You can enable exactly-once delivery on pull subscriptions, use ordering keys when order matters, and configure dead-letter topics for messages that keep failing.
The rest of this is premium
Written by an OurLogic creator – 60% of your payment goes to them.
Log in to unlockFeedback & comments (1)
Pooja Shah
Great system-design answer. Would love a Part 3 on Vertex AI pipelines.