In this article, let us have a brief overview of Data ingestion and Data Modeling in Salesforce Data cloud. In Salesforce Data Cloud, data is first ingested from the source and stored in a data lake object. Data is retrieved from the source using connectors, which establish communication between servers, allowing continuous data access. Data Streams support these connectors by dictating the frequency and timing of these connections. Data modeling involves mapping data streams to a unified data model, creating a harmonized view across various data sources.
Data ingestion
Connectors and data streams play a crucial role in ensuring seamless data access. Let’s explore the process of retrieving data from various sources using connectors and data streams:
Data Sets:
- When you bring in data from a source, it’s essential to preserve its original structure so that if you make a mistake or need to adjust requirements, you can always go back to the original data shape.
- Include all fields from the data set without modification.
- Extend the data set by creating additional fields.
- We can use formulas to clean up nomenclature or standardize data.
Connectors:
- Connectors establish communication between servers and data sources (e.g., databases, APIs, files).
- They enable data retrieval by maintaining a connection to the source.
- Connectors ensure continuous access to the data.
Data Streams:
- Each data set is represented by a data stream in Data Cloud.
- Data streams define how often and when data connections are established.
- They dictate the frequency of data retrieval.
- For example, a real-time data stream might fetch updates every minute, while a batch stream could run daily.

- Data Source Selection:
Choose between a previously connected data source or authenticate a new one (such as cloud storage).

- Starting Bundle or Object Selection:
Select the starting bundle or specify the object or filename you want to work with.

- Source Configuration:
Within Data Cloud, there is a place to write in the name of the source. Next to the source, you specify the data set that you’re bringing in from that source by filling out the Object Label and Object API Name.
- Field Verification and Primary Key:
Refer to the primary key of the data set in your matrix and designate that field as the primary key when defining the data source object.

- Enhance Data (Optional):
Add formula fields to cleanse or derive new data and create additional fields as needed.
- Data Refresh Settings:
Choose how often you want your data to be updated. Set the refresh schedule.
Data modeling
To enable interaction between data sets, they must conform to a universal language. This is achieved in the data modeling phase using Salesforce’s Customer 360 Data Model. This model includes various objects covering multiple subject areas and is extensible, allowing the addition of custom attributes to standard objects and the creation of new custom objects. These custom objects can be defined based on how they relate to existing objects. The Customer 360 Data Model assigns a semantic context to source objects, creating a harmonized data layer abstracted from the underlying source objects. This standardization ensures that data, regardless of its origin, is consistent and usable across different tasks.
In this article we have covered an overview of Data ingestion and Data Modeling in Salesforce Data cloud. To know more, click Data Cloud Features and Learning Path (salesforce.com)
Take Five Consulting is a technology company, based in Virginia U.S., that specializes in the Mortgage Banking vertical especially LOS implementation and application development. Take Five Consulting creates and implement mortgage technology and software specifically for Mortgage Industry.


