ThatQuiz Test Library Take this test now
formattive-2(DataSciece)
Contributed by: Dhanalakshmi
  • 1. Decision tree is the most powerful for ____
A) prediction
B) classification
C) None of these
D) both a and b
  • 2. Decision trees can handle_____
A) High dimensional data
B) medium dimensional data
C) None of these
D) low diamesional data
  • 3. In Decision-tree algorithm At the beginning, we consider the whole training set as ____
A) leaf
B) steam
C) None of these
D) root
  • 4. ___is the measure of uncertainty of a random variable, it characterizes the impurity of an arbitrary collection of examples.
A) Information Gain
B) Entropy
C) Gini Index
D) None of these
  • 5. What are the advantages of the decision tree?
A) Non-linear patterns in the data can be captured easily
B) Both
C) What are the advantages of the decision tree?
D) None of these
  • 6. Which of the following is correct with respect to random forest?
A) Random forest are difficult to interpret but very less accurate
B) forest are Random difficult to interpret but often very accurate
C) None of these
D) Random forest are easy to interpret but often very accurate
  • 7. Which of the following is an essential process in which the intelligent methods are applied to extract data patterns?
A) Data Mining
B) Data Selection
C) Warehousing
D) Text Mining
  • 8. What is KDD in data mining?
A) Knowledge data house
B) Knowledge Data definition
C) Knowledge Discovery Database
D) Knowledge Discovery Data
  • 9. For what purpose, the analysis tools pre-compute the summaries of the huge amount of data?
A) In order to maintain consistency
B) To obtain the queries response
C) For data access
D) For authentication
  • 10. What are the functions of Data Mining?
A) Cluster analysis and Evolution analysis
B) Prediction and characterization
C) All of the above
D) Association and correctional analysis classification
  • 11. Which one of the following statements about the K-means clustering is incorrect?
A) All of the above
B) K-means clustering can be defined as the method of quantization
C) The goal of the k-means clustering is to partition (n) observation into (k) clusters
D) The nearest neighbor is the same as the K-means
  • 12. In data mining, how many categories of functions are included?
A) 3
B) 5
C) 2
D) 4
  • 13. What is the importance of using PCA before the clustering? Choose the most complete answer
A) Find the explained variance
B) Find which dimension of data maximize the features variance
C) Avoid bad features
D) Find good features to improve your clustering score
  • 14. Following the steps to run a PCA's algorithm, why is so important standardize your data?
A) Find the features which can best predicts Y
B) Make the training time more fast
C) Use Standardize the best practices of data wrangling
D) data allows other people understand better your work
  • 15. . Which of the following model model include a backwards elimination feature selection routine?
A) MCV
B) MCRS
C) All of the mentioned
D) MARS
  • 16. Which of the following function is a wrapper for different lattice plots to visualize the data?
A) plotsample
B) None of the mentioned
C) featurePlot
D) levelplot
  • 17. Which of the following can be used to impute data sets based only on information in the training set?
A) preProcess
B) All of the above
C) process
D) postProcess
  • 18. The function preProcess estimates the required parameters for each operation.
A) True
B) False
  • 19. Which of the following can also be used to find new variables that are linear combinations of the original set with independent components?
A) None of the mentioned
B) PCA
C) ICA
D) SCA
  • 20. . The preProcess class can be used for many operations on predictors.
A) True
B) False
Created with That Quiz — where a math practice test is always one click away.