> For the complete documentation index, see [llms.txt](https://awsarch.adot8.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://awsarch.adot8.com/module-15.md).

# Module 15

Data engineering patterns

The five Vs of data characteristics are value, veracity, volume, velocity, and variety. Each of these characteristics impact decision-making with data

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2FHpc65ewfzSGIebvMFL1O%2Fimage.png?alt=media&amp;token=2914eb1a-8e91-4bd3-8ae3-dadc03b4650b" alt=""><figcaption></figcaption></figure>

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2F7ph6zIlLg2lAmdMI4qB1%2Fimage.png?alt=media&amp;token=71974ee5-a01b-4199-977e-799d70cf57e1" alt=""><figcaption></figcaption></figure>

#### Three-pronged strategy to build data infrastructure

* Modernize
  * Move to a cloud-based infrastructure and purpose-built services to reduce undifferentiated lifting.
* Unify
  * Create a single source of truth for data, and make the data available across the organization.
* Innovate
  * Apply artificial intelligence and machine learning (AI/ML) to find new insights in the data

{% hint style="info" %}
Data lake = Raw/Unstructured data

Data Warehouse = Structured data
{% endhint %}

### Elements of a data pipeline

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2Fdp2E5pq2Qar4w1gwQ6lo%2Fimage.png?alt=media&amp;token=e2b038ef-56e5-46c6-8380-37ae2ae624ef" alt=""><figcaption></figcaption></figure>

#### Homogeneous ingestion pattern

Essentially just extracting the data and storing it in the same format it was orginally made in

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2F9UX5sx2U8MNDFlbkfFy9%2Fimage.png?alt=media&amp;token=37beff54-60cc-4b85-b96c-50d877855f97" alt=""><figcaption></figcaption></figure>

#### Heterogeneous ingestion patterns

| Extract, transform and load (ETL)                                                      | Extract, load, and transform (ELT)                                           |
| -------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- |
| Works well with structured data that is destined for a data warehouse                  | Works well for unstructured data that is destined for a data lake            |
| Stores data that is ready to be analyzed, so this pattern can save time for an analyst | Offers flexibility to create new queries (analysts can access more raw data) |

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2FOvPNpzy2XdBMzi8UjpLk%2Fimage.png?alt=media&amp;token=cf4b76d7-368e-4020-8f3f-f19d0f9e162e" alt=""><figcaption></figcaption></figure>

#### Batch and Streaming processing patterns

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2F3kbgksMIjxNZvNzJTUZC%2Fimage.png?alt=media&amp;token=5e5535e1-72e3-48b9-afe9-cd8fcc33fe86" alt=""><figcaption></figcaption></figure>

### AWS tools to ingest data

### Amazon App Flow

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2FdHKgZCZfYDvSqdYF6I0p%2Fimage.png?alt=media&amp;token=37ba5de5-275e-470a-8b34-4dd0680496ed" alt=""><figcaption></figcaption></figure>

* Provides the ability to transfer data between SaaS applications and AWS services
* Offers reuse of available service integrations with available Amazon AppFlowAPIs
* Provides data transformation capabilities

#### AWS DataSync

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2FcskpHClCX6hOZQdKQ76k%2Fimage.png?alt=media&amp;token=d15ab8c2-6b46-48c0-84ae-9d46b8256287" alt=""><figcaption></figcaption></figure>

* Is a fully managed data migration service•Simplifies, automates, and accelerates copying file and object data to and from AWS storage services
* Is optimized for speed•Includes encryption and integrity validation
* Preserves metadata when moving data

#### AWS Data Exchange

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2FzX0NIB1fXFDGpG6OncYY%2Fimage.png?alt=media&amp;token=abf2c175-c5fd-4754-8a98-a3be8ea0ddc9" alt=""><figcaption></figcaption></figure>

* Provides customers with a way to find, subscribe to, and use third-party data in the cloud
* Bridges the gap between providers and subscribers who exchange data by supporting data delivery through files, tables, and APIs
* Simplifies finding, preparing, and using data in the cloud

### Processing Data in AWS

#### Batch ingestion and processing

Example:

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2FGZE25rCLxkm7qsEgcjOA%2Fimage.png?alt=media&amp;token=3d4d84d1-b7be-4846-b88b-2039d26b2102" alt=""><figcaption></figcaption></figure>

Batch processing is for:

* For reporting purposes
* When dealing with large datasets
* When the analytics use case is  focused more on aggregating or transforming data and less on real-time analysis of data

#### AWS Glue

* Used for batch processing
* Is a data integration service that helps automate and perform ETL tasks as part of ingesting data into a pipeline
* Provides the ability to read and write data from multiple systems and databases
* Simplifies batch and streaming ingestion

#### AWS Glue Components

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2FKojIHi4ltd6YkaOiPnb6%2Fimage.png?alt=media&amp;token=e02726bc-d887-4285-99df-755c0c21d03b" alt=""><figcaption></figcaption></figure>

#### Example

<figure><img src="https://3899036363-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FsmRwCapKRPjknpALqp9Y%2Fuploads%2Fek16ohqTQOsMPH1Tsig2%2Fimage.png?alt=media&amp;token=143173e1-6557-4fe2-afd8-51a885b50b80" alt=""><figcaption></figcaption></figure>

#### AWS Glue Transformation types

| .csv                                                                            | .parquet                            | Convert .csv to .parquet                       |
| ------------------------------------------------------------------------------- | ----------------------------------- | ---------------------------------------------- |
| Is the most common format to store tabular data                                 | Stores data in a columnar fashion   | Speeds up analytics workloads                  |
| Isn’t efficient to store or manipulate large amounts of data (more than 15 GBs) | is optimized for storage            | Over time, saves storage space, cost, and time |
|                                                                                 | Is suitable for parallel processing |                                                |

###
