HISTORY OF DATA WAREHOUSING
  • Data warehouses extend the transformation of data into information
  • In the 1990’s executives became less concerned with the day-to-day business operations and more concerned with overall business functions
  • The data warehouse provided the ability to support decision making without disrupting the day-to-day operations

DATA WAREHOUSE FUNDAMENTALS
  • Data warehouse – a logical collection of information – gathered from many different operational databases – that supports business analysis activities and decision-making tasks
  • The primary purpose of a data warehouse is to aggregate information throughout an organization into a single repository for decision-making purposes
  • The primary difference between a database and a data warehouse is that a database stores information for a single application, whereas a data warehouse stores information from multiple databases, or multiple applications, and external information such as industry information 
  • This enables cross-functional analysis, industry analysis, market analysis, etc., all from a single repository
  • Data warehouses support only analytical processing (OLAP)
  • Extraction, transformation, and loading (ETL) – a process that extracts information from internal and external databases, transforms the information using a common set of enterprise definitions, and loads the information into a data warehouse
  • The ETL process gathers data from the internal and external databases and passes it to the data warehouse
  • The ETL process also gathers data from the data warehouse and passes it to the data marts
Data mart – contains a subset of data warehouse information


  • The data warehouse modeled in the above figure compiles information from internal databases or transactional/operational databases and external databases through ETL
  • It then send subsets of information to the data marts through the ETL process

MULTIDIMENSIONAL ANALYSIS AND DATA MININ
  • Databases contain information in a series of two-dimensional tables
  • In a data warehouse and data mart, information is multidimensional, it contains layers of columns and rows
  • Dimension – a particular attribute of information
  • Each layer in a data warehouse or data mart represents information according to an additional dimension
  • Dimensions could include such things as:
    • Products
    • Promotions
    • Stores
    • Category
    • Region
    • Stock price
    • Date
    • Time
    • Weather
  • Why is the ability to look at information based on different dimensions critical to a business success?
    • Ans:  The ability to look at information from different dimensions can add tremendous business insight
    • By slicing-and-dicing the information a business can uncover great unexpected insights
  • Cube – common term for the representation of multidimensional information



  • Users can slice and dice the cube to drill down into the information
  • Cube A represents store information (the layers), product information (the rows), and promotion information (the columns)
  • Cube B represents a slice of information displaying promotion II for all products at all storesCube C represents a slice of information displaying promotion III for product B at store 2
  • Data mining – the process of analyzing data to extract information not offered by the raw data alone
  • Data mining can begin at a summary information level (coarse granularity) and progress through increasing levels of detail (drilling down), or the reverse (drilling up)To perform data mining users need data-mining tools
  • Data-mining tool – uses a variety of techniques to find patterns and relationships in large volumes of information and infers rules that predict future behavior and guide decision making
  • Data-mining tools include query tools, reporting tools, multidimensional analysis tools, statistical tools, and intelligent agents

INFORMATION CLEANSING OR SCRUBBING
          An organization must maintain high-quality data in the data warehouse
          Information cleansing or scrubbing – a process that weeds out and fixes or discards inconsistent, incorrect, or incomplete information
          Contact information in an operational system


Taking a look at customer information highlights why information cleansing and scrubbing is necessary
Customer information exists in several operational systems
In each system all details of this customer information could change form the customer ID to contact information
Determining which contact information is accurate and correct for this customer depends on the business process that is being executed


          Standardizing Customer name from Operational Systems



          Information cleansing activities


          Accurate and complete information


          Why do you think most businesses cannot achieve 100% accurate and complete information?
          If they had to choose a percentage for acceptable information what would it be and why?
§  Some companies are willing to go as low as 20% complete just to find business intelligence
§  Few organizations will go below 50% accurate – the information is useless if it is not accurate
          Achieving perfect information is almost impossible
§  The more complete and accurate an organization wants to get its information, the more it costs
§  The tradeoff between perfect information lies in accuracy verses completeness
§  Accurate information means it is correct, while complete information means there are no blanks
§  Most organizations determine a percentage high enough to make good decisions at a reasonable cost, such as 85% accurate and 65% complete

BUSINESS INTELLIGENCE
BI is information that people use to support their decision-making efforts
Principle BI enablers include:
          Technology
          Even the smallest company with BI software can do sophisticated analyses today that were unavailable to the largest organizations a generation ago. The largest companies today can create enterprisewide BI systems that compute and monitor metrics on virtually every variable important for managing the company. How is this possible? The answer is technology—the most significant enabler of business intelligence.
          People
          Understanding the role of people in BI allows organizations to systematically create insight and turn these insights into actions. Organizations can improve their decision making by having the right people making the decisions. This usually means a manager who is in the field and close to the customer rather than an analyst rich in data but poor in experience. In recent years “business intelligence for the masses” has been an important trend, and many organizations have made great strides in providing sophisticated yet simple analytical tools and information to a much larger user population than previously possible.
          Culture
          A key responsibility of executives is to shape and manage corporate culture. The extent to which the BI attitude flourishes in an organization depends in large part on the organization’s culture. Perhaps the most important step an organization can take to encourage BI is to measure the performance of the organization against a set of key indicators. The actions of publishing what the organization thinks are the most important indicators, measuring these indicators, and analyzing the results to guide improvement display a strong commitment to BI throughout the organization.

The End of Chapter 8 : Accessing Organizational Information – Data Warehouse by syahirahzfri. 
Thank you for reading :)



RELATIONAL DATABASE FUNDAMENTALS
Information is everywhere in an organization and information is stored in databases.
  • Database – maintains information about various types of objects (inventory), events (transactions), people (employees), and places (warehouses)
  • Database models include:
    • Hierarchical database model – information is organized into a tree-like structure (using parent/child relationships) in such a way that it cannot have too many relationships
    • Network database model – a flexible way of representing objects and their relationships
    • Relational database model – stores information in the form of logically related two-dimensional tables
Entities and Attributes
Entity – a person, place, thing, transaction, or event about which information is stored
Attributes (fields, columns) – characteristics or properties of an entity class

Keys and Relationships
  • Primary key – a field (or group of fields) that uniquely identifies a given entity in a table
  • Foreign key – a primary key of one table that appears an attribute in another table and acts to provide a logical relationship among the two tables
RELATIONAL DATABASE ADVANTAGES
Database advantages from a business perspective include
  • Increased flexibility
  • Increased scalability and performance
  • Reduced information redundancy
  • Increased information integrity (quality)
  • Increased information security
Increased Flexibility
A well-designed database should:
  • Handle changes quickly and easily
  • Provide users with different views
  • Have only one physical view
    • Physical view – deals with the physical storage of information on a storage device
  • Have multiple logical views
    • Logical view – focuses on how users logically access information

Increased Scalability and Performance
A database must scale to meet increased demand,  while maintaining acceptable performance levels
  •  Scalability – refers to how well a system can adapt to increased demands
  • Performance – measures how quickly a system performs a certain process or transaction

Reduced Information Redundancy
  • One of the primary goals of a database is to eliminate information redundancy by recording each piece of information in only one place
  • Databases reduce information redundancy
    • Redundancy – the duplication of information or storing the same information in multiple places
  • Inconsistency is one of the primary problems with redundant information

Increase Information Integrity (Quality)
  • Information integrity – measures the quality of information
  • Integrity constraint – rules that help ensure the quality of information

Increased Information Security
  • Information is an organizational asset and must be protected
  • Databases offer several security features including:
    • Password – provides authentication of the user
    • Access level – determines who has access to the different types of information
    • Access control – determines types of user access, such as read-only access
DATABASE MANAGEMENT SYSTEMS
software through which users and application programs interact with a database

DATA-DRIVEN WEB SITES
A data-driven Web site is an interactive Web site kept constantly updated and relevant to the needs of its customers through the use of a database. Data-driven Web sites are especially useful when the site offers a great deal of information, products, or services. Web site visitors are frequently angered if they are buried under an avalanche of information when searching a Web site. A data-driven Web site invites visitors to select and view what they are interested in by inserting a query, which the Web site then analyzes and custom builds a Web page in real-time that satisfies the query. The figure displays a Wikipedia user querying business intelligence and the database sending back the appropriate Web page that satisfies the user’s request.

Data-Driven Web Site Business Advantages
  • Development: Allows the Web site owner to make changes any time—all without having to rely on a developer or knowing HTML programming. A well-structured, data-driven Web site enables updating with little or no training.
  • Content management: A static Web site requires a programmer to make updates. This adds an unnecessary layer between the business and its Web content, which can lead to misunderstandings and slow turnarounds for desired changes.
  • Future expandability: Having a data-driven Web site enables the site to grow faster than would be possible with a static site.  Changing the layout, displays, and functionality of the site (adding more features and sections) is easier with a data-driven solution.
  • Minimizing human error: Even the most competent programmer charged with the task of maintaining many pages will overlook things and make mistakes. This will lead to bugs and inconsistencies that can be time consuming and expensive to track down and fix. Unfortunately, users who come across these bugs will likely become irritated and may leave the site. A well-designed, data-driven Web site will have ”error trapping” mechanisms to ensure that required information is filled out correctly and that content is entered and displayed in its correct format.
  • Cutting production and update costs: A data-driven Web site can be updated and ”published” by any competent data entry or administrative person. In addition to being convenient and more affordable, changes and updates will take a fraction of the time that they would with a static site. While training a competent programmer can take months or even years, training a data entry person can be done in 30 to 60 minutes.
  • More efficient:  By their very nature, computers are excellent at keeping volumes of information intact. With a data-driven solution, the system keeps track of the templates, so users do not have to. Global changes to layout, navigation, or site structure would need to be programmed only once, in one place, and the site itself will take care of propagating those changes to the appropriate pages and areas. A data-driven infrastructure will improve the reliability and stability of a Web site, while greatly reducing the chance of ”breaking” some part of the site when adding new areas.
  • Improved Stability: Any programmer who has to update a Web site from ”static” templates must be very organized to keep track of all the source files. If a programmer leaves unexpectedly, it could involve re-creating existing work if those source files cannot be found. Plus, if there were any changes to the templates, the new programmer must be careful to use only the latest version. With a data-driven Web site, there is peace of mind, knowing the content is never lost—even if your programmer is.
Integrating Information among Multiple Databases
Integration – allows separate systems to communicate directly with each other
  • Forward integration – takes information entered into a given system and sends it automatically to all downstream systems and processes
  • Backward integration – takes information entered into a given system and sends it automatically to all upstream systems and processes
  • One of the biggest benefits of integration is that organizations only have to enter information into the systems once and it is automatically sent to all of the other systems throughout the organization
  • This feature alone creates huge advantages for organizations because it reduces information redundancy and ensures accuracy and completeness
  • Without integrations an organization would have to enter information into every single system that requires the information from marketing and sales to billing and customer service
Integrating Information among Multiple Databases
Building a central repository specifically for integrated information
  • The above figure displays an example of customer information integrated using this method
  • Users can create, read, update, and delete in the main customer repository, and it is automatically sent to all of the other databases
  • This method does not follow the business process when building the integrations
  • Business-critical integrity constraints still need to be built to ensure information is only ever entered into the customer repository, otherwise the information will become out-of-sync

The End of Chapter 7: Storing Organizational Information by syahirahzfri. 
Thank you for reading :)