flink-query-senior-data-engineerWorld-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure. Expertise in Python,...
Install via ClawdBot CLI:
clawdbot install wu-uk/flink-query-senior-data-engineerGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Potentially destructive shell commands in tool definitions
rm -rf /Calls external URL not in known-safe list
http://spark-history:18080Audited Apr 18, 2026 · audit v1.0
Generated May 7, 2026
Migrate a legacy on-premise data warehouse to a modern cloud-based lakehouse architecture using Apache Spark and Delta Lake. This involves redesigning ETL pipelines, implementing incremental loads, and ensuring data quality across the transition.
Build a streaming pipeline using Kafka and Flink to process financial transactions in real-time, detecting fraud patterns with sub-second latency. Includes monitoring consumer lag and schema drift to maintain high throughput and reliability.
Design a batch pipeline that aggregates daily sales, inventory, and customer behavior data from multiple sources into Snowflake. Use dbt for transformations and Airflow for orchestration, serving a Tableau dashboard for business intelligence.
Implement a data quality and governance framework for a healthcare provider, ensuring HIPAA compliance. Develop validation rules for patient records, automate lineage tracking, and set up alerts for data anomalies.
Create a streaming architecture to ingest and process millions of IoT sensor messages per second using Kinesis and Spark Structured Streaming. Include monitoring for data freshness and dead letter queues to handle malformed data.
Offer end-to-end data pipeline design, implementation, and maintenance to clients as a subscription service. Generate recurring revenue through monthly retainers for pipeline monitoring and optimization.
Provide fixed-fee consulting engagements for data architecture design, migration, and training. Revenue from project-based contracts for specific deliverables like pipeline setup or performance tuning.
Develop and license a data quality validation framework that can be integrated into existing pipelines. Charge per deployment or on a per-data-asset basis.
💬 Integration Tip
Leverage the provided scripts (e.g., pipeline_orchestrator.py) as templates; customize YAML configs to match existing infrastructure such as Airflow connections or schema registry endpoints.
Scored May 7, 2026
Use Redis effectively for caching, queues, and data structures with proper expiration and persistence.
MarkItDown is a Python utility from Microsoft for converting various files (PDF, Word, Excel, PPTX, Images, Audio) to Markdown. Useful for extracting structu...
Database Designer - POWERFUL Tier Skill
Execute SQL queries, inspect Snowflake databases, schemas, tables, views, warehouses, and data pipeline resources. Use this skill when users want to query da...
Skill Oracle — Curated documentation of quality ClawHub skills. Markdown tables telling agents which tools work and which are empty. Not an API or code library.
Database Schema Designer