To request a post on a specific topic or if you have any questions email James@StatisticsSolutions.com
Showing posts with label cluster analysis. Show all posts
Showing posts with label cluster analysis. Show all posts

Thursday, August 6, 2009

Cluster Analysis

Statistics Solutions is the country's leader in statistical consulting and cluster analysis. Contact Statistics Solutions today for a free 30-minute consultation.

Cluster analysis can be used in market research problems for satisfying certain purposes. This document will discuss the utilization of cluster analysis.

Cluster analysis can be used to segment consumers on the basis of allowances sought from the purchase of the product. Cluster analysis can be used to identify the homogeneous group of buyers in the market.

Cluster analysis is that kind of technique that is used to discover or understand the structures within a complex body of structures. In other words, cluster analysis gets involved in the segmentation of data. There are three methods or approaches in cluster analysis to serve this purpose: the hierarchical methods, the partitioning methods and the two step procedure.

The hierarchical methods in cluster analysis wrap up the data row-wise. On the other hand, the partitioning method (or the non hierarchical method) in cluster analysis wraps the data into a specified number of segments and further interchanges the variables to improve the measure of effectiveness in the data. Finally, the two step procedure in cluster analysis determines the perfect number of clusters by comparing the values of model choice criteria across different clustering solutions.

The procedure of performing cluster analysis involves formulating the problem, selecting a measure of distance, selecting a procedure for clustering, determining the number of clusters, interpreting the clusters profile and assessing the validity of clustering.

With every analysis, a researcher becomes familiar with different kinds of variables. In the case of cluster analysis, the researcher becomes familiar with variables based on past research. An adaptable measure of distance is selected in cluster analysis. One can use a commonly used measure called Euclidean distance in cluster analysis.

There is a lower triangle matrix called similarity (or the distance coefficient matrix) in cluster analysis. This consists of pair-wise distances between the objects or cases.

The cluster results in cluster analysis are indicated with the help of a graphical display device called a dendrogram. The desirable clusters in cluster analysis are the ones which are widely separated and are explicit.

The researcher working upon cluster analysis should always be aware that in cluster analysis, no clustering solution will be accepted without some assessment of the cluster analysis’s reliability and validity. The procedures used in assessing the reliability of clustering results in cluster analysis are quite complex. These procedures tell the researcher to perform cluster analysis on the same data using some different measures of distance. Then, the researcher compares the results across the measures in cluster analysis in order to determine the stability or the validity of the cluster solutions. The final procedure tells the researcher to perform different methods or approaches of clustering in cluster analysis and then to compare all the procedures or approaches simultaneously. By doing this procedure, the researcher can assess the validity of the cluster results in cluster analysis.

Additionally, a procedure of deleting variables randomly in cluster analysis is done by many researchers. The researcher then performs a clustering based on the reduced set of variables in cluster analysis. Then, a comparison on the result of the ones with original variables and the ones with random variables is done in cluster analysis.

Friday, June 26, 2009

Cluster Analysis

The term cluster analysis was first used by Tryon in 1939. Cluster analysis is a multivariate technique, which is used for segmentation in research. In cluster analysis, a cluster is a group of similar objects. In marketing research, cluster analysis is used for segmentation of similar objects, which is similar in buying habits, demographic characteristics, or psychographics.

Statistics Solutions can assist with cluster analysis and additional statistical analysis. Contact Statistics today for a free 30-minute consultation.

Cluster analysis is also known as the data reduction technique. In other words, cluster analysis is an exploratory data analysis technique that groups similar objects in such a way that the distance between the two objects is minimal, or the group of similar objects is grouped in such a way that the association between the variables is maximized. Cluster analysis seeks to minimize the within group variance and maximize the between group variance. For example, it can be used if an A FMCG Company wants to match the profile of the target audience in terms of lifestyle, attitude and perception. In this case, the marketing manager prepares a questionnaire with 20 statements and then performs a cluster analysis on the 20 statements in three clusters. These days, researchers have developed similar techniques for cluster analysis, which have different names, like numerical taxonomy, Q-analysis, typology analysis, classification analysis, etc. To perform cluster analysis, there are a number of techniques available that are based on the procedure, which is used to measure the distance and clustering algorithm.

Assumptions in cluster analysis:

1. The sample taken for cluster analysis should be representative of the population.
2. In cluster analysis, it is assumed that multiple collinearity is minimal.
3. In cluster analysis, the outlier affects the results. Thus, cluster analysis assumes that there is an absence of outliers.
4. In cluster analysis, data may be metric, non-metric, or a combination of both.
5. Naturally occurring groups must be present in the data.
The process of cluster analysis:
1. A cluster analysis starts with the N×K database.
2. In the second step of cluster analysis, different steps are used to create the N×N matrix. In this N×N matrix, each case is similar or dissimilar to another case, based on the k number of variables.
3. In the third step, by using different algorithms, subjects are sorted in to statistically significant groups. In this group, subjects are as homogenous or different to each other as possible.

In cluster analysis, the N×N matrix is created by using one of the following methods: Squared Euclidean Distance, Pearson correlation coefficient, Cosine of vector variables, Minkowski metric, Mahalanobis D2, City block or Manhattan distances, Jaccard’s coefficient, Chebychev distance metric, Gower’s coefficient, etc. In SPSS, most of the techniques are available.

Clustering algorithms: In cluster analysis cluster algorithms are of two types:

1. Hierarchical methods: Cluster analysis in hierarchical method algorithms involve single average (or nearest neighbor), complete average (or furthest neighbor), average linkage, centroid methods, etc.
2. Non-hierarchical methods: A Cluster analysis non-hierarchical algorithm involves sequential threshold methods, parallel methods, optimization methods, etc.

Determining the number of clusters in the data: In cluster analysis, there is no particular procedure that is used to determine the cluster in the data. The following procedure is used in cluster analysis to determine the cluster in the data:

1. Clustering coefficient: In cluster analysis, the coefficient size shows the homogeneity of the objects being merged.
2. Dendrogram: Dendrogram is the pictorial representation of the cluster in the data. In dendrogram, we can see how the observations are combined in each cluster.
3. Vertical icicle: Vertical icicle is another pictorial way to find the number of clusters in the data. In vertical icicle, blanks are clusters and X’s indicate the number of members per cluster.

Cluster analysis and SPSS: Cluster analysis can be conducted using SPSS. To conduct cluster analysis in SPSS, click on the “Analysis” option and select “classify option” and select “required cluster analysis.”