Big data is a concept I discuss regularly in my posts. Data is collected about nearly everything we do, because that information is valuable. Think about your organization: how many data sets or reports does your organization rely on? Whether your organization maintains hundreds of data assets or hundreds of thousands, chances are there's a lot of data that has to be analyzed before it can be used.
With SAS Data Governance (formerly SAS Information Catalog), your organization can take advantage of easy-to-use tools to discover, classify, track, and manage data assets across cloud and on-premises environments. In this post, I'll show you how to use the new data profiling agent to discover and profile your assets. Keep reading to see how to gain insights on a lot of data with little work!
Data Profiling Agents
Several types of agents are available in SAS Data Governance, including schema extraction agents, data profiling agents, discovery agents, and flow metadata agents. For an overview of each type of agent and their purpose, visit the documentation. Starting with the SAS Viya 2026.06 stable release, the new data profiling agent replaces the discovery agent, which was used in earlier releases of SAS Information Catalog (now SAS Data Governance) for profiling data assets.
The data profiling agent offers more detailed and customizable profiling capabilities to enhance the metadata discovered by a schema extraction agent. You can select specific profiler metrics to include in the data analysis, optimize your agent, and set filter rules to include or exclude assets. For a complete overview of the profilers and options available for data profiling agents, visit the documentation.
Scenario
In SAS Viya, I've created the ORACAS caslib which connects to an Oracle database. Though I can see the list of 18 total tables in ORACAS in SAS Data Explorer, I don't know any information about them. I'll use SAS Data Governance to fill these gaps in the metadata.
Select any image to see a larger version.
Mobile users: To view the images, select the "Full" version at the bottom of the page.
Running Schema Extraction Agents
First, I have to be logged into SAS Viya as an administrator. Only administrators can maintain data agents in SAS Data Governance.
After opening SAS Data Governance, I'll go to the Agents tab. Here, I can manage, run, and monitor my agents. Additionally, I can browse the available agents by connections or by all agents.
When I expand the ORACAS connection, I can see a schema extraction agent. A schema extraction agent is created for each data source by default to index their assets and retrieve their basic metadata. The data profiling agent depends on the metadata collected by a schema extraction agent, so I need to run this agent first. For more information on schema extraction agents, visit the documentation.
The schema extraction agent runs quickly and discovers 18 new assets. I can click the number of new assets for review.
This opens a list of all tables analyzed by the agent. By default, the properties of the first table in the list are shown in the right panel.
If I open any table, I can see the information found by the schema extraction agent. In this example, I'll open the SALESSTAFF table. I can see basic column metadata on the Overview tab.
However, if I select the Column Profile tab, key metrics like minimum, median, mean, and maximum values for each column are missing. This is where the data profiling agent will come in. At the top of the screen, I can select the Kick off analysis option to profile this one table, but I'll set up an agent to profile all tables in the library instead.
Creating and Running Data Profiling Agents
I'll return to the Agents tab in SAS Data Governance, select the ORACAS caslib (beneath the ORACAS connection name), and select + Agent to add the data profiling agent.
First, I'll set the properties. I'll name the agent Data profiling agent ORACAS and provide a brief description.
Next, I'll select which profiling agents I want to run. The key metrics agents (Numeric Data Analysis and Character Data Analysis) are selected by default. Additionally, you can select advanced analytics agents (Correlation Analysis and Textual Summarization) and a data governance agent (Semantic Classification). The best agents to choose will depend on the data you are profiling. Since my data is fairly basic, I'll select the key metrics and data governance agents.
For any agents that are profiling text columns, you also should specify the country and language to define the locale of the data (for example, English in the United States or Spanish in Mexico).
The remaining steps are optional. You can specify optimization settings, such as overriding the default performance settings, running the agent only when the data changes, adding custom analysis options, or adjusting the retention settings for the profiling history. I'll add in the option crawlPersist=1, which ensures that the agent crawls persistent tables in CAS. To see all valid analysis options, visit the documentation.
You can also choose to set filter rules. When selecting the option to process Tables that match filter rules, you're able to define filters and whether to include or exclude the tables that meet those rules. For this post, I'll process all tables.
Lastly, you can set scheduling options to run the agent at a specific time and frequency. You can add triggers like with any other scheduling capabilities in SAS Viya.
Once I'm finished with configuration, I'll save the agent. Now, it appears under the list of agents for ORACAS, and I can run it.
The data profiling agent will take longer to run than the schema extraction agent, because it's doing a more thorough analysis of each table in the library. In this case, my data profiling agent ran in about two and a half minutes. I can see that all 18 assets were updated, and I can click the number of assets for review.
I'll open the same table I reviewed last time, SALESSTAFF, and see how the table information has changed. The Overview tab now includes an analysis summary and I can see that the overview is more detailed, including information like data classification and information privacy.
Additionally, the Column Profile tab is now filled out with the results from the data profiling agent. Now, I can see key metrics for any of the columns in the table, including distinct value counts, common data metrics, a frequency distribution analysis, and a box plot.
Note that the metrics will vary for character type columns. For example, instead of minimum and maximum values, metrics include minimum and maximum string lengths. Instead of a box plot for the data, a chart displaying the results of pattern analysis is displayed.
Thanks to the data profiling agent, I now have detailed information about every table and column in ORACAS. I can re-run this agent anytime to retrieve the newest metrics, or I can update my agent to include more profilers and retrieve even more metrics.
Summary
In this post, I introduced the new data profiling agent in SAS Data Governance (formerly SAS Information Catalog), which is a flexible, simple-to-use tool for analyzing your data assets. For more information about the newest features in SAS Data Governance, check out the 2026.06 stable release notes.
Have you used discovery agents for profiling data previously? Are there specific features of the new data profiling agent that will improve your data governance approaches? Share your thoughts, questions, and feedback below!
Find more articles from SAS Global Enablement and Learning here.
Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.
The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.