“It was the best of times, it was the worst of times, it was the age of wisdom, it was the age of foolishness, it was the epoch of belief, it was the epoch of incredulity, it was the season of Light, it was the season of Darkness, it was the spring of hope, it was the winter of despair, we had everything before us, we had nothing before us,……...“
In any discussion involving Inmon & Kimball, I often find the debate becomes abstract and 'hand-wavy', with few concrete worked examples. As a result, the real conceptual differences are obscured and become more about one approach being 'better' than the other. Much of the confusion stems from unstated assumptions on both sides, which are rarely surfaced or challenged, even though they are central to understanding when each approach is appropriate.
For example, the claim that Kimball architectures involve multiple redundant 'interfaces' from data marts to source systems. Taken literally, this makes little sense; any competent modern architecture would ingest data from each source once into a shared landing area, right?. Took me while to figure out that the real premise was not multiple interfaces/ingestions, but multiple integrations; in early Kimball implementations, each data mart may independently resolve entities, apply business rules, resolve encoding, and reconcile sources. Integration logic is therefore duplicated across marts, whereas in an Inmon style architecture that logic is centralised and executed once, with downstream marts inheriting a consistent enterprise meaning.
Both sides, however, rely on assumptions that are often left implicit. Kimball advocates may assume conformed dimensions, shared staging layers, and strong governance; Inmon advocates may assume central authority, stable definitions, and organisational tolerance for slower initial delivery. These conditions do not hold universally.
A more productive discussion would move away from 'Inmon vs Kimball' and towards example-led elaborative analysis of where integration logic lives, how much divergence is acceptable, and what trade-offs are appropriate for a given organisational context - because, as architects we know, no single solution is best for every scenario!
"For example, the claim that Kimball architectures involve multiple redundant 'interfaces' from data marts to source systems. Taken literally, this makes little sense; any competent modern architecture would ingest data from each source once into a shared landing area, right?."
I am guessing you were not around in the early 90s.
In the early 90s we did not have incremental extract from operational systems. We had to take full dumps or period dumps. These were down to tapes. You could not afford the disk to create a shared landing area in an RDBMS. The "landing area" was tapes.
So we wrote cobol code reading those tapes. Metaphor Computer Systems, the company Ralph Co Founded in 1982 wrote exactly what is described. Multiple interfaces to those tapes to send data into dimensional models.
I can confirm this because I did the performance tuning on the Standard and Chartered Bank dimensional data warehouse built by Metaphor in 1993. In September 1993 I attended the Metaphor Users Conference in San Francisco. I was introduced to their Professional Services Manager and to the man who took over from Ralph in 1989 when Ralph left after the IBM acquisition.
That mans first name is Steve and he confirmed that all the customers were doing the same thing S&C were doing.
Bill was the Key Note Speaker at that conference and he was invited by Coca Cola, worlds largest users of Metaphor at the time.
Bill gave his pitch about Time Variance and Stability Analysis. I was a 10 year IT veteran and and IBM Systems Engineer at the time and I did not understand a word Bill said. But everyone else in the room was frantically taking notes so I did too! LOL!
After his presentation I waited in the line to introduce myself to Bill and ask him if he had written any books. He told me he had written Building the Data Warehouse so I got a pal in Metaphor to buy it and send it to me. I STILL did not understand it and in 1994 we brought Bill out to Australia to take him fishing. He caught a 50kg Yellow Fin Tuna.
When you take a man deep sea tuna fishing off the coast of Sydney and he has the luck to land the biggest fish of the day?! You become pals! LOL! And that's what happened.
In was not until the seminar we ran with Bill speaking that I actually "got" that he was talking about ARCHIVING ALL THE DATA IN THE COMPANY.
I still remember putting up my hand and saying: "But Bill, that is too much data, we can't possibly archive ALL the data in a company."
And Bill reply was: "Peter, we are talking about volumes of data beyond human comprehension today."
You could have heard a pin drop in the room and there were LOTS of people standing.
That was the first time it dawned on MANY of us in the room that what Bill was proposing was to store ALL the data in the company for perhaps as long as 10 years.
This was a staggering idea for the time. So Steve and I discussed this and in 1995, as far as I am aware, I was the first man to EVER build an incrementally updated landing area, that fed data into a staging area, that then fed data into a dimensional model.
That was the IBM RB2020 Banking System Data Warehouse project I did in 1995. It was the first time Change Data Capture had been used by anyone I knew in the data warehouse environment because CDC was brand new on DB2.
The reason we could do it was that we wanted to sell more IBM hardware. The overall Banking System project was later cancelled and that implementation was never put into production.
Then in 1996 Hitachi Data Systems hired me to launch their data warehousing practice where we sold Sequent machines. The Global data warehouse practice manager at Sequent was an old buddy of mine from Metaphor and they were selling dimensional models. Their hardware was perfect for dimensional models.
Steve, this guy, and myself put our heads together to talk about what Hitachi might do and between us we decided that we would have a landing area, a staging area, and be committed to dimensional models. I sold the first two such projects in February 1997 but politics got in the way, I sold a third one in July 1997 and we had it implemented by the end of the year. That, as far as I know, was the first time ever an "all your data" data warehouse was built with a landing area, a staging area, and a dimensional model and had made it to production.
Meanwhile, also in 1996, Steve had an insurance client where an archive would be great but many of the requirements also required a dimensional model. We went back and forth for a week or two and some how we came up with "let's do both". Stever was then working for PwC and running the PwC Data Modeling area of their data warehouse practice.
This was one of the BIG insurance companies in the US.
So Steve designed an Archive like Bill proposes and a dimensional model on top like Ralph proposes and they rolled that out. Early results were back in by the end of 1996 and Steve reported it worked really well but was VERY expensive which PwC liked.
So PwC standardised on archival models plus dimensional models over the top where they could get the customer to pay for it. I then moved to PwC in Australia in 1998 to run the data warehouse practice for Australia and New Zealand. I did sell the largest ever data warehouse deal in Australia in 1998 and even then we did not sell the archive because we could not do it in the time frames. We sold landing, staging, dimensional models and what was delivered was landing to dimensional models and they had all the problems associated with doing that.
So what Bill has written is exactly correct. It is exactly what was happening. And I know because I was there. Ok?
Well I'm glad I made that comment... so I could hear your fascinating story!
My point mainly is that many such arguments due to technology limitations of 90's do not hold water today, as of course, many of todays limitations will be long gotten over in another 20 years time. Here's wishing you a amazing new year 2026! and hope to hear more such interesting anecdotes - this is exactly the kind of practical detail that is missing. Thanks for your reply. Regards.
Hi Datastrat, I will be releasing a post today that goes into the details of the evolution of data models.
I invented the then worlds most productive ETL tool in 1995. It has evolved over the years to where now it's possible to map 1,000 fields from source to target in one very long working day.
From 1995 to 2017 the mapping rate was 1,000 fields per 220 hour work month. So you can see to cram that into a 15 hour day opens up all sorts of possibilities.
I sold a LOT of mega projects at 1,000 fields per 220 hour work month. We are now 80 percent cheaper.
I also invented what I call Mega Models. This is where many customers data can reside in one data warehouse and only one set of ETL is needed. And even when the many customers data is in the one data warehouse image I have invented a way where no custers data co habit in one table.
This removes pretty much all existing technical limitations. We know what data warehousing will look like in 20 years because I have invented it.
That people are still using GUI based tools to develop ETL is an embarrassment to our profession.
for some reason I do not know Margy has never talked to me. With the release of Sybase IWS pretty much everything Ralph was talking about was superseded. However, because we were selling the models for USD150K what was in them was Confidential
I did share what we had done with archiving data inside the dimensional model with both Bill and Ralph in 2003. In 2002 I invented the one more step needed to perfectly archive data in a dimensional model. Even the worlds best data models paid me the extremely high compliment of "that's a good idea" at breakfast. LOL!
Once I invented that "one last step"? We knew we were going to sell and absolute TRUCK LOAD of these models because archiving data in the dimensional model itself reduces costs 50% with zero loss of functionality.
We instantly became 50% cheaper than the next competitor with zero loss of function. Until I made that invention we had the same issue as Ralph did. We did not archive data perfectly in the dimensional models and if an archive was needed then we have to create a separate archive.
Anyway, Sybase made the decision to keep what we had confidential and both Bill and Ralph respected my request to keep what I showed them confidential.
Sean Kelly and I went on to sell two copies of our new telco models before I was woke cancelled. I am still woke cancelled and our industry has suffered quite a bit for not taking my advice.
The whole area of business intelligence is regarded as a bit of a joke now on the business side of the house. Deservedly so.
Maybe, now that I can have a substack? People in our industry segment will be able to hear what I have to say and they can start to improve their standing with the business side of the house. But I am not holding my breath.
it is just setting the record straight. I have been so sick and tired of this "argument" over the last 30 years. People have treated it like a religious argument.
There was never any "conflict" between Bill and Ralph. I mean I met Bill at the 1993 Metaphor Users Conference where he was invited to be the key note speaker by Coca Cola, largest users of Metaphor at the time.
In 1997 the worlds best data modeler figured out a way to archive data in dimensional models. I did his class in February 2001 and immediately realised what he had done. He worked for Sybase at the time and they were called the Sybase Industry Warehouse Studio Models and they remain the worlds best way to develop models.
The key breakthrough was the ability to properly archive data inside the dimensional models. Sybase was selling them for USD150K per copy and they were selling like hot cakes.
Sybase was bought by SAP who had SAP BW and the IWS Models died on the vine. Sean Kelly and I asked SAP to allow us to create a new generation of models. SAPs position was we could go back to what was sold to Sybase in 1999 and begin from there and not compete with BW. So we did and we sold two telco models to Carphone Warehouse and Skytalk in the UK before I was doxed and "woke cancelled" in 2010.
The Sybase IWS models were a LONG way in front of what Ralph was publishing which is why we could sell them for USD150K at Sybase and EUR100K later on with Sean Kelly.
I now give away copies of these models at the base level. They are still the best models available and now the base levels are free.
Just so you know, both Bill and Ralph are personal friends of mine and both of them helped me make a couple million dollars income for which I am most grateful.
Where can I read more about the stage 1 of Kimball? In the Toolkit Book I know it always sounded pretty much that you must try to integrate domains and conform dimensions. So it was interesting (and surprising) to hear about the 1st stage which indeed seems to be very redundant and far away from SVOT
you can just read Ralphs first book if you want to go all the way back. Today the best way to implement an enterprise data warehouse for high volume transaction based businesses via the style of dimensional models I am giving away. We used to sell them for EUR100K but I have been woke cancelled and now I am giving them away.
Great writeup. My team is working with some banks now to upgrade some 1990's stuff and it is not straight forward. The Cobol guys are all retiring and so much is not transitioned. Its not like you can just convert Cobol to Python.
Even today, what Ralph and Margy published has not reached the level that the Sybase Industry Warehouse Studio models did by 2002.
"I worked on a project a few years back where we tried the fast dimensional route and hit exactly the integration wall you describe around year two, had to basicaly rebuild everything."
Yep. Out there in the world today there are LOTS of people building dimensional data warehouses that will be crap over time all because I have been woke cancelled since 2010 and people do not want to use our data models because they don't want to "offend women" by talking to me. Pretty stupid really.
In any discussion involving Inmon & Kimball, I often find the debate becomes abstract and 'hand-wavy', with few concrete worked examples. As a result, the real conceptual differences are obscured and become more about one approach being 'better' than the other. Much of the confusion stems from unstated assumptions on both sides, which are rarely surfaced or challenged, even though they are central to understanding when each approach is appropriate.
For example, the claim that Kimball architectures involve multiple redundant 'interfaces' from data marts to source systems. Taken literally, this makes little sense; any competent modern architecture would ingest data from each source once into a shared landing area, right?. Took me while to figure out that the real premise was not multiple interfaces/ingestions, but multiple integrations; in early Kimball implementations, each data mart may independently resolve entities, apply business rules, resolve encoding, and reconcile sources. Integration logic is therefore duplicated across marts, whereas in an Inmon style architecture that logic is centralised and executed once, with downstream marts inheriting a consistent enterprise meaning.
Both sides, however, rely on assumptions that are often left implicit. Kimball advocates may assume conformed dimensions, shared staging layers, and strong governance; Inmon advocates may assume central authority, stable definitions, and organisational tolerance for slower initial delivery. These conditions do not hold universally.
A more productive discussion would move away from 'Inmon vs Kimball' and towards example-led elaborative analysis of where integration logic lives, how much divergence is acceptable, and what trade-offs are appropriate for a given organisational context - because, as architects we know, no single solution is best for every scenario!
Hi DataStrat,
"For example, the claim that Kimball architectures involve multiple redundant 'interfaces' from data marts to source systems. Taken literally, this makes little sense; any competent modern architecture would ingest data from each source once into a shared landing area, right?."
I am guessing you were not around in the early 90s.
In the early 90s we did not have incremental extract from operational systems. We had to take full dumps or period dumps. These were down to tapes. You could not afford the disk to create a shared landing area in an RDBMS. The "landing area" was tapes.
So we wrote cobol code reading those tapes. Metaphor Computer Systems, the company Ralph Co Founded in 1982 wrote exactly what is described. Multiple interfaces to those tapes to send data into dimensional models.
I can confirm this because I did the performance tuning on the Standard and Chartered Bank dimensional data warehouse built by Metaphor in 1993. In September 1993 I attended the Metaphor Users Conference in San Francisco. I was introduced to their Professional Services Manager and to the man who took over from Ralph in 1989 when Ralph left after the IBM acquisition.
That mans first name is Steve and he confirmed that all the customers were doing the same thing S&C were doing.
Bill was the Key Note Speaker at that conference and he was invited by Coca Cola, worlds largest users of Metaphor at the time.
Bill gave his pitch about Time Variance and Stability Analysis. I was a 10 year IT veteran and and IBM Systems Engineer at the time and I did not understand a word Bill said. But everyone else in the room was frantically taking notes so I did too! LOL!
After his presentation I waited in the line to introduce myself to Bill and ask him if he had written any books. He told me he had written Building the Data Warehouse so I got a pal in Metaphor to buy it and send it to me. I STILL did not understand it and in 1994 we brought Bill out to Australia to take him fishing. He caught a 50kg Yellow Fin Tuna.
When you take a man deep sea tuna fishing off the coast of Sydney and he has the luck to land the biggest fish of the day?! You become pals! LOL! And that's what happened.
In was not until the seminar we ran with Bill speaking that I actually "got" that he was talking about ARCHIVING ALL THE DATA IN THE COMPANY.
I still remember putting up my hand and saying: "But Bill, that is too much data, we can't possibly archive ALL the data in a company."
And Bill reply was: "Peter, we are talking about volumes of data beyond human comprehension today."
You could have heard a pin drop in the room and there were LOTS of people standing.
That was the first time it dawned on MANY of us in the room that what Bill was proposing was to store ALL the data in the company for perhaps as long as 10 years.
This was a staggering idea for the time. So Steve and I discussed this and in 1995, as far as I am aware, I was the first man to EVER build an incrementally updated landing area, that fed data into a staging area, that then fed data into a dimensional model.
That was the IBM RB2020 Banking System Data Warehouse project I did in 1995. It was the first time Change Data Capture had been used by anyone I knew in the data warehouse environment because CDC was brand new on DB2.
The reason we could do it was that we wanted to sell more IBM hardware. The overall Banking System project was later cancelled and that implementation was never put into production.
Then in 1996 Hitachi Data Systems hired me to launch their data warehousing practice where we sold Sequent machines. The Global data warehouse practice manager at Sequent was an old buddy of mine from Metaphor and they were selling dimensional models. Their hardware was perfect for dimensional models.
Steve, this guy, and myself put our heads together to talk about what Hitachi might do and between us we decided that we would have a landing area, a staging area, and be committed to dimensional models. I sold the first two such projects in February 1997 but politics got in the way, I sold a third one in July 1997 and we had it implemented by the end of the year. That, as far as I know, was the first time ever an "all your data" data warehouse was built with a landing area, a staging area, and a dimensional model and had made it to production.
Meanwhile, also in 1996, Steve had an insurance client where an archive would be great but many of the requirements also required a dimensional model. We went back and forth for a week or two and some how we came up with "let's do both". Stever was then working for PwC and running the PwC Data Modeling area of their data warehouse practice.
This was one of the BIG insurance companies in the US.
So Steve designed an Archive like Bill proposes and a dimensional model on top like Ralph proposes and they rolled that out. Early results were back in by the end of 1996 and Steve reported it worked really well but was VERY expensive which PwC liked.
So PwC standardised on archival models plus dimensional models over the top where they could get the customer to pay for it. I then moved to PwC in Australia in 1998 to run the data warehouse practice for Australia and New Zealand. I did sell the largest ever data warehouse deal in Australia in 1998 and even then we did not sell the archive because we could not do it in the time frames. We sold landing, staging, dimensional models and what was delivered was landing to dimensional models and they had all the problems associated with doing that.
So what Bill has written is exactly correct. It is exactly what was happening. And I know because I was there. Ok?
Well I'm glad I made that comment... so I could hear your fascinating story!
My point mainly is that many such arguments due to technology limitations of 90's do not hold water today, as of course, many of todays limitations will be long gotten over in another 20 years time. Here's wishing you a amazing new year 2026! and hope to hear more such interesting anecdotes - this is exactly the kind of practical detail that is missing. Thanks for your reply. Regards.
Hi Datastrat,
here you are. I am sure you will like this if you liked my comment above.
https://peterandrewnolan.substack.com/p/ibi-071-the-history-of-data-warehouse
Hi Datastrat, I will be releasing a post today that goes into the details of the evolution of data models.
I invented the then worlds most productive ETL tool in 1995. It has evolved over the years to where now it's possible to map 1,000 fields from source to target in one very long working day.
From 1995 to 2017 the mapping rate was 1,000 fields per 220 hour work month. So you can see to cram that into a 15 hour day opens up all sorts of possibilities.
I sold a LOT of mega projects at 1,000 fields per 220 hour work month. We are now 80 percent cheaper.
I also invented what I call Mega Models. This is where many customers data can reside in one data warehouse and only one set of ETL is needed. And even when the many customers data is in the one data warehouse image I have invented a way where no custers data co habit in one table.
This removes pretty much all existing technical limitations. We know what data warehousing will look like in 20 years because I have invented it.
That people are still using GUI based tools to develop ETL is an embarrassment to our profession.
https://decisionworks.com/2004/03/differences-of-opinion is article from 2004 where the diagram is an example of what NOT to do. Ridiculously unfair argument from Inmon.
Hi Donald,
for some reason I do not know Margy has never talked to me. With the release of Sybase IWS pretty much everything Ralph was talking about was superseded. However, because we were selling the models for USD150K what was in them was Confidential
I did share what we had done with archiving data inside the dimensional model with both Bill and Ralph in 2003. In 2002 I invented the one more step needed to perfectly archive data in a dimensional model. Even the worlds best data models paid me the extremely high compliment of "that's a good idea" at breakfast. LOL!
Once I invented that "one last step"? We knew we were going to sell and absolute TRUCK LOAD of these models because archiving data in the dimensional model itself reduces costs 50% with zero loss of functionality.
We instantly became 50% cheaper than the next competitor with zero loss of function. Until I made that invention we had the same issue as Ralph did. We did not archive data perfectly in the dimensional models and if an archive was needed then we have to create a separate archive.
Anyway, Sybase made the decision to keep what we had confidential and both Bill and Ralph respected my request to keep what I showed them confidential.
Sean Kelly and I went on to sell two copies of our new telco models before I was woke cancelled. I am still woke cancelled and our industry has suffered quite a bit for not taking my advice.
The whole area of business intelligence is regarded as a bit of a joke now on the business side of the house. Deservedly so.
Maybe, now that I can have a substack? People in our industry segment will be able to hear what I have to say and they can start to improve their standing with the business side of the house. But I am not holding my breath.
Isn’t this a tired old repost of his article from decades ago? If you want to understand Kimball and Ross, read Kimball and Ross, not this dross.
Hi Donald,
it is just setting the record straight. I have been so sick and tired of this "argument" over the last 30 years. People have treated it like a religious argument.
There was never any "conflict" between Bill and Ralph. I mean I met Bill at the 1993 Metaphor Users Conference where he was invited to be the key note speaker by Coca Cola, largest users of Metaphor at the time.
In 1997 the worlds best data modeler figured out a way to archive data in dimensional models. I did his class in February 2001 and immediately realised what he had done. He worked for Sybase at the time and they were called the Sybase Industry Warehouse Studio Models and they remain the worlds best way to develop models.
The key breakthrough was the ability to properly archive data inside the dimensional models. Sybase was selling them for USD150K per copy and they were selling like hot cakes.
Sybase was bought by SAP who had SAP BW and the IWS Models died on the vine. Sean Kelly and I asked SAP to allow us to create a new generation of models. SAPs position was we could go back to what was sold to Sybase in 1999 and begin from there and not compete with BW. So we did and we sold two telco models to Carphone Warehouse and Skytalk in the UK before I was doxed and "woke cancelled" in 2010.
The Sybase IWS models were a LONG way in front of what Ralph was publishing which is why we could sell them for USD150K at Sybase and EUR100K later on with Sean Kelly.
I now give away copies of these models at the base level. They are still the best models available and now the base levels are free.
Just so you know, both Bill and Ralph are personal friends of mine and both of them helped me make a couple million dollars income for which I am most grateful.
Thanks a lot for the post!
Where can I read more about the stage 1 of Kimball? In the Toolkit Book I know it always sounded pretty much that you must try to integrate domains and conform dimensions. So it was interesting (and surprising) to hear about the 1st stage which indeed seems to be very redundant and far away from SVOT
Hi Tim,
you can just read Ralphs first book if you want to go all the way back. Today the best way to implement an enterprise data warehouse for high volume transaction based businesses via the style of dimensional models I am giving away. We used to sell them for EUR100K but I have been woke cancelled and now I am giving them away.
Marvelous piece when seen and combined in other contexts.
Those 4 levels:
- very current
- current
- less than current
- older
are a reflection of what happens in organisations processing the flows.
Pondering this in light of Data Mesh approaches. Any articles around the contrasts would be greatly appreciated.
Great writeup. My team is working with some banks now to upgrade some 1990's stuff and it is not straight forward. The Cobol guys are all retiring and so much is not transitioned. Its not like you can just convert Cobol to Python.
Even today, what Ralph and Margy published has not reached the level that the Sybase Industry Warehouse Studio models did by 2002.
"I worked on a project a few years back where we tried the fast dimensional route and hit exactly the integration wall you describe around year two, had to basicaly rebuild everything."
Yep. Out there in the world today there are LOTS of people building dimensional data warehouses that will be crap over time all because I have been woke cancelled since 2010 and people do not want to use our data models because they don't want to "offend women" by talking to me. Pretty stupid really.