DATA MANAGEMENT IN THE AGE OF AI
How has AI changed data management in the corporation
DATA MANAGEMENT IN THE AGE OF AI
By W H Inmon
With contributions from Jamie Knowles, Jamie Knowles @semanticsarethesolution
Data management has travelled a long road from the beginning of the computer industry to the day and age of generative AI. Data management has been elevated from the role of simple technician to the role of great strategic importance in the world of generative AI.
That path has had many steps and missteps along the way. But today data management arrives at the position of being crucial to the success and adoption of generative AI in the corporation. If AI does not operate on believable data, then any analysis done by AI is questionable.
IN THE BEGINNING…
In the very beginning, systems were built by application developers. The developer had many tasks – requirements gathering, methodology conformation, programming, testing and the like. One of the tasks the early application developer had was the task of defining the data on which the system would run.
The task of designing the data and the structure of the data for the application system being built was but one of the many tasks of development.
TRANSACTION PROCESSING APPLICATIONS
The world progressed from early simple systems to transaction processing systems. With transaction processing systems came a whole new world of response time, system availability, transaction execution integrity and so forth. What was needed was a data base designer who could incorporate all of these needs into the data base design.
With transaction applications, the life of the data administrator became quite a bit more complex.
THE FLOOD OF TRANSACTION PROCESSING APPLICATIONS
And soon there was a plethora of transaction processing systems – a flood. It was recognized that there was a need for control and coordination across the many transaction processing data bases that were being built.
What was needed was a data base administrator.
DATA INTEGRITY
And with the growth of the many data bases came the realization that data integrity was an issue. In the spiders web of applications systems that resulted, the same element of data appeared in many places, each with a different value in each place. No longer did the organization have the problem of having data. Now the organization had the problem of believing what data existed. In one place the data showed one value. I another place the data showed another value. And resolving this discrepancy was no small task.
Furthermore, adding more application systems and more technology only exaggerated and accelerated the problem of lack of data integrity
.A technical solution was not needed. What was needed was an architectural solution.
AN ARCHITECTURAL SOLUTION
The architectural solution that resulted was the separation of operational data from analytical data. Thus born was the data warehouse.
THE NEED FOR A DATA MODEL
In order to achieve the architectural solution, the world discovered that it need a data model. The data model was an abstraction of the types of data that existed and the data that needed to exist in the future.
The role of the CIO – chief information officer - appeared.
The world of operational systems and data warehouses was a world that was exclusively a structured world. In a structured world each record format is identical in a data base. Only the contents of the record differ from each other. All applications, data bases, and transaction processing data bases were based on structured data. The structured data was managed by a data base management system – a dbms.
The world of data management changed dramatically with the movement to data warehousing.
And with the tremendous influx of new transaction processing systems came the recognition that data management was an important function. With transaction processing systems came a new role for systems. With transaction processing systems, when the system was malfunctioning, the business and its customers were directly and negatively affected. ATM machines did not work. Airline reservation systems would not function. It became very obvious to the corporation at the business level that reliable, accurate systems were needed, and those systems ran on data.
The enhanced role of the importance of transaction processing systems elevated the value of data administration and management to the organization.
THE ELEMENTS OF DATA MANAGEMENT IN THE STRUCTURED ENVIRONMENT
So what were the elements of data administration in the day of the chief data officer?
One function was that of managing the data model. Not only did the data model need to be built, but the data model needed to be maintained and altered as businesses changed. And every business was subject to constant change. There were changes in economic conditions. In competition. In new technologies. In legislation and market conditions. The business of every corporation changes over time.
And as these changes to business occurred, the data model needed to be kept abreast of those changes.
A related task of the structured data manager is the task of conducting data base design. Data base design – in its infancy - started by conforming to the rules of normalization. However, the performance requirements of transaction processing mandated that some amount of denormalization of data be done in order to achieve the performance objectives of the transaction processing system..
The data manager is responsible for the balancing of the goals and needs for normalization and denormalization of data.
Another task of the data manager in the structured environment is that of constructing the DDL that defines the data base. The DDL is patterned after the data model.
Capacity planning is another task of the data manager. Capacity planning is especially important in transaction processing. The organization had a need to know when the limits of a machine were going to be reached.
Monitoring both data base activity and data base growth is another ongoing activity of data management in the structured environment.
THE ROLE OF THE DATA MODEL
The data model is a necessary component of the transformation of application data to analytical data, or the data warehouse.
DATA MANAGEMENT IN THE STRUCTURED ENVIRONMENT
In many respects data management in the structured environment is very much a hands on experience. When the data manager in the structured environment sees something that needs an amendment, the data manager goes directly to the problem and fixes it.
ENTER UNSTRUCTURED/TEXTUAL DATA
As interesting and as important as structured data is in the corporation, structured data is hardly the only kind of data found in the corporation. Another pervasive data type is unstructured or textual data. It is estimated that as much as 90% of the data in the corporation is textual data. Furthermore, there is much very important data that is wrapped up in the form of text.
Structured data is entirely different from unstructured/textual data.
And textual data – like structured data – needs to be managed by the data management organization
.THE TEXTUAL DATA MODEL
There is no classical data model that fits well with the textual environment. Instead, the equivalent of the classical data model is the ontology/taxonomy. Stated differently, the ontology/taxonomy model fits the textual environment just as the classical data model fits the structured environment.
However analogically similar they are, the classical structured data model is VERY different from the ontology/taxonomy.
A typical mistake made by data management when they first engage with the world of text is to try to apply the skills needed for classical data modelling and data management in the structured world to the world of text.
Such an misapplication of skills is a serious misfit and yields very few actual positive results
.DATA MANAGEMENT IN THE GENERATIVE AI ENVIRONMENT
So what are the tasks for data management in the textual environment? The first and most important task is to ensure that raw text before it enters the LLM is vetted and that extraneous, non business relevant text does not find its way into the LLM. If the data manager does nothing else, the insurance that no extraneous text ever enters the LLM is sufficient.
The removal of extraneous, non business related data from the stream of data entering the LLM is important because of several reasons –
1) Removing extraneous text from the stream of text entering the LLM makes the LLM much more efficient to process against. The savings are in hard core dollars that are saved every time a query is issued.
2) Removing extraneous, non business related text provides focus for the LLM. It is clear what the LLM represents.
THE ELDM – ENTERPRISE LOGICAL DATA MODEL
Another task of the data manager for the generative AI environment is that of marrying the ontology/taxonomy environment with the larger corporate ELDM. The ELDM – enterprise logical data model – is necessary to meld the classical data model with the ontology/taxonomy.
In addition the data manager of the generative AI environment needs to keep the ontology/taxonomy data model up to date, as business needs change.
BROAD SWATHS OF DATA
Yet another task for the data manager for the generative AI environment is that of the selection of broad swaths of raw text from sources such as the Internet. This selection process sets the stage for the secondary process of selecting business relevant data from the text presented to the LLM.
ONGOING MAINTENANCE
The data manager for the generative AI environment has got to do constant surveillance and management of the ontology/taxonomy environment. Business and world conditions are constantly changing. The ontology/taxonomy needs to be kept in synch with these changing conditions.
INDIRECT CONTROL OF THE LLM
Data management in the generative AI environment is much like the manipulation of a puppet. The data manager in the generative AI environment does not have the direct control that the data manager in the structured environment has. Instead, the data manager in the generative AI environment has indirect control. The indirect control is the management of what text is fed the LLM.
The following figure shows the difference between data management in the structured environment and data management in the generative AI environment. One type of control is direct. The other type of control is indirect.
THE DATA MANAGER IN THE ORGANIZATION CHART
In the beginning data management was merely one more activity that the system developer had to tend to. There was no specific corporate function for data management in the beginning.
When transaction processing became a reality, it was recognized that data base design was its own activity (that of course had to integrate with other activities of development.) Thus evolved the role of the data base designer. The data base designer was the first formal recognition of the corporation of the need for data management.
Soon there were LOTS of data bases, as the spiders web environment grew. From the spiders web environment grew the recognition of the need for a data base administrator. There were simply too many data bases and too many elements of data to be managed. A larger corporate role was recognized.
Soon the confusion over the disintegrity of data in the corporation that was found in the spiders web environment led to the need for the chief information officer, or the CIO.
And finally the need to integrate textual data with structured data to have a true corporate understanding of data led to the CDO – chief data officer.
With the advent of the many forms of data, there grew great confusion and great diversity of opportunities and challenges with data. The job of managing grew both in complexity and in importance.
And the corporation recognized that need.
THE CORPORATE PROGRESSION
The progression of the role of the data administration is shown in the following chart.
CENTRALIZATION
One of issues of data management that has radically changed in the transition from early systems to the world of generative AI is the whole idea of centralization of data. In order to achieve a cohesive enterprise understanding of data and its meaning, it is mandatory that there be a centralized understanding of data.
But the whole notion of what is meant by centralization has also changed.
In the very early days of computing centralization referred to the physical centralization of data. Data bases were physically centralized into large mainframe computers.
But as time passed, physical centralization became an impossibility.
Data was on personal computers. Data was on the Internet. Data was on corporate systems. Data was in fact – everywhere.
And with the diaspora of data, came an even greater need for centralization. But the impetus for centralization took place not at the physical level but at the semantic level.
What someone was referring to on the Internet needed to be synchronized with what was being said in the corporation, for example.
The result was that centralization took place at the semantic level, not the physical level.
The vehicle for shepherding the semantics of data became the ELDM – enterprise logical data model.
The ELDM held the enterprise understanding of data centrally, where the data resided in many physical locations.
Bill Inmon lives in Denver with his wife and his two Scotty dogs – Lena and Rollie. Lena got to go to the art fair in Evergreen last weekend. She came home and told Rollie all about it. He was jealous. But he will get to go on the next adventure up into the mountains. Then Lena will be jealous.
LLM MGMT, LLC – a Bill Inmon company - is a company for the preparation of raw text to enter the corporate LLM. LLM MGMT separates business related text from extraneous text before the text is entered into the LLM. LLM MGMT allows text to be gathered automatically from a wide variety of sources in order for the text to be entered into the corporate LLM. LLM MGMT allows structured data bases to be read and the data that is gathered is entered into the corporate LLM.
For more information, see the web site – llmmgmt.com.


































@bill inmon is no @elon musk. But, he's no @ed wolfe either!! I L-O-V-E the way Bill explains things for *ALL* of us (even ABOVE elon!).
Thank you