using-data-engineering-agent-skills — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited using-data-engineering-agent-skills (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Start here before changing code, SQL, orchestration, contracts, or infrastructure. This skill maps the user request to the right data engineering workflow so the agent does not skip specification, quality, governance, replay, or rollback thinking.
Do not stop here once the task has been classified. Load the actual execution skill after triage.
data-specificationpipeline-planning-and-task-breakdownSFTP, or partner-managed feeds: use file-and-partner-feed-ingestionGlue Data Catalog, Lake Formation, or AWS-native catalog governance: use glue-data-catalog-and-lake-formation-governancePython implementation work: use python-data-engineering-and-pipeline-packagingScala JVM data jobs: use scala-data-engineering-on-jvm-runtimesJava connectors or data services: use java-data-engineering-and-integration-servicesUnity Catalog or Databricks lakehouse governance: use unity-catalog-and-lakehouse-governancePurview or Azure-native governance: use microsoft-purview-and-azure-data-governanceDataplex, policy tags, or BigQuery governance: use dataplex-and-bigquery-governanceMySQL versus NoSQL or operational-store choice: use operational-datastore-selection-relational-and-nosqlETL, ELT, or transformation-boundary redesign: use etl-elt-and-modernization-strategytest-data-preparation-and-synthetic-datalower-environment-data-masking-and-obfuscationdata-quality-and-contract-testingdata-resiliency-testing-and-failure-injectiondata-platform-disaster-recovery-and-business-continuityincident-triage-and-pipeline-recoverysafe-backfill-and-replay-orchestration, orchestration-and-backfills, and data-migration-and-platform-cutoverspark-serverless-reliability-and-state-managementkafka-resilience-and-schema-evolutionmcp-data-observability-integrationlineage-pii-and-governancePII, PCI, HIPAA, PHI, or audit-bound data handling: use data-security-compliance-and-regulated-dataregional-data-compliance-and-sovereigntyESG reporting data products: use esg-and-sustainability-regulatory-reportingschema-evolution-and-contract-migrationsInformatica, Talend, or legacy ETL modernization: use enterprise-etl-and-data-integration-modernizationmainframe-modernization-and-data-offloaddata-platform-operating-model-and-service-ownershipdata-quality-platforms-and-rule-managementpython-data-engineering-and-pipeline-packagingscala-data-engineering-on-jvm-runtimesjava-data-engineering-and-integration-serviceswarehouse-and-schema-design + dbt-and-analytics-engineeringsnowflake-native-pipelines-and-governancebigquery-and-dataform-platform-engineeringAWS, Azure, GCP, or Databricks: use the matching governance skill plus the matching platform presetspark-and-distributed-processingairflow-and-workflow-orchestrationstreaming-and-messaging-systemscdc-and-incremental-loading or debezium-and-kafka-connect-cdclakehouse-table-format-engineering or delta-lake-and-medallion-architectureterraform-and-data-platform-infrastructureregional-data-compliance-and-sovereignty or esg-and-sustainability-regulatory-reportingenterprise-etl-and-data-integration-modernizationaws-data-engineeringazure-data-engineeringgcp-data-engineeringdatabricks-lakehouse-engineeringsnowflake-modern-data-platforminformatica-data-integrationtalend-data-integrationmulti-cloud-hybrid-data-engineeringtemplates/source-contract.yaml or templates/dataset-contract.yamltemplates/metric-contract.yamltemplates/incident-runbook.md, templates/backfill-plan.yaml, or templates/release-gate-evidence.yamltemplates/schema-change-plan.yamlstarter-packs/aws-lakehouse-starter.yamlstarter-packs/databricks-medallion-starter.yamlstarter-packs/warehouse-analytics-starter.yamlstarter-packs/streaming-reliability-starter.yamlstarter-packs/production-reliability-starter.yamlstarter-packs/data-platform-cicd-release-starter.yamlstarter-packs/resiliency-testing-starter.yamlstarter-packs/validation-security-review-starter.yamlstarter-packs/privacy-governance-starter.yamlstarter-packs/regional-compliance-and-esg-reporting-starter.yamlstarter-packs/test-data-lower-environments-starter.yamlstarter-packs/enterprise-etl-modernization-starter.yamlPrefer runnable scaffolds when the team needs executable proof; use architecture blueprints for /spec and /plan only. See examples/README.md for the full type column.
Runnable:
examples/dbt-warehouse-marts/examples/databricks-delta-medallion/examples/aws-s3-glue-athena-iceberg/examples/kafka-flink-streaming/examples/aws-serverless-spark-msk-reliability/Blueprint:
examples/api-saas-to-warehouse-ingestion/examples/privacy-retention-deletion-workflow/examples/multi-cloud-warehouse-cutover//spec/plan/build/validate/review/backfill/shipreferences/api-saas-ingestion-checklist.mdreferences/file-ingestion-checklist.mdreferences/data-quality-checklist.mdreferences/security-compliance-regulated-data-checklist.mdregistry/assets.jsonreferences/data-validation-and-testcase-patterns.mdreferences/data-resiliency-testing-patterns.mdreferences/data-platform-dr-bcp-checklist.mdreferences/data-platform-operating-model-checklist.mdreferences/platform-native-governance-patterns.mdreferences/data-platform-security-checklist.mdreferences/data-engineering-anti-patterns.mdreferences/etl-elt-modernization-checklist.mdreferences/mainframe-modernization-checklist.mdreferences/progressive-data-release-patterns.mdreferences/test-data-preparation-checklist.mdreferences/lower-environment-masking-checklist.mdreferences/enterprise-etl-modernization-checklist.mdreferences/regional-compliance-and-data-sovereignty-checklist.mdreferences/esg-and-sustainability-reporting-checklist.mdreferences/streaming-architecture-patterns.mdreferences/spark-serverless-reliability-patterns.mdreferences/kafka-production-guardrails.mdreferences/mcp-data-observability-patterns.mdreferences/cloud-data-engineering-architecture-patterns.mdreferences/pipeline-orchestration-patterns.mdreferences/data-quality-tooling-and-rule-management.mdreferences/README.mdreferences/streaming-checklist.mdreferences/incident-recovery-checklist.mdreferences/schema-migration-checklist.mdreferences/observability-and-sla-checklist.md| Rationalization | Reality |
|---|---|
| "It is only a small pipeline tweak." | Small data changes still break dashboards, contracts, or downstream jobs. Classify the work correctly first. |
| "We already know the stack." | Stack knowledge does not replace contract, replay, quality, or publish decisions. |
| "We can skip the example because the work is custom." | Examples reduce guessing and make rollout patterns safer even when the domain differs. |
| "Backfill is just rerunning the job." | Replay work can amplify bad data and downstream impact if windowing and reconciliation are unclear. |
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.