Memory hook: Locate the failing layer, then optimize the measured bottleneck.
Must remember
Monitor ingestion, transformation, semantic-model refresh and consumption. Track job state, duration, freshness, completeness, rejected records and capacity utilization. Alerts need owners and actionable thresholds. Capacity contention can slow several unrelated items at once, while a single malformed record may fail only one pipeline.
For pipeline/Dataflow failures, inspect the failing activity/step, parameters, credentials, gateway/connection, schema and source throttling. For notebooks, examine execution logs, library/environment versions, resource exhaustion and data paths. Eventstream/Eventhouse failures may involve source connections, schemas, transformation expressions or ingestion policies. T-SQL errors need actual statement/schema/permission evidence; shortcut errors need target/connection checks.
Optimize Lakehouse tables by addressing small files, partition design and supported compaction/layout options. V-Order targets suitable columnar read efficiency; it is not a universal remedy for every write-heavy workload. Cleanup such as VACUUM affects historical-file availability, so retention and readers matter. Do not delete underlying files manually to solve a table error.
Tune Spark partitions, join strategies, skew, caching and executor resources from observed stages. Tune warehouse queries by reducing unnecessary scans and expensive joins and checking concurrency. For streaming, examine throughput, batch/window/state behavior and backpressure. Pipelines benefit from appropriate parallelism, not unlimited simultaneous source requests.
Diagnosis drill: a daily dashboard is stale despite a green copy activity. Check transformation completion, target partitions, semantic-model refresh and report access in order. End-to-end correctness crosses more than one successful task.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Many jobs suddenly slower | Inspect shared capacity and concurrency. |
| One Spark stage dominates | Inspect skew, shuffle and partitioning before adding resources. |
| Lakehouse has many tiny files | Evaluate supported compaction and write-pattern changes. |
Traps
- More parallelism can overwhelm a source or capacity.
- A green upstream task does not prove a fresh downstream report.
Active recall
1. What is backpressure?
Downstream processing cannot keep up with incoming work.
2. Why avoid manual Delta file deletion?
It can break transaction-log consistency and readers.
3. Why can caching hurt?
It consumes memory and may not benefit a reused/appropriate workload.
4. What should a refresh investigation include?
Source credentials, data changes, processing success and downstream model/report state.
5. What should precede optimization?
A measured bottleneck and a correctness baseline.