ThatQuiz Test Library Take this test now
COELE1-2
Contributed by: Billo
  • 1. "Mini-batch Gradient Descent" is often preferred because it:
A) Is guaranteed to converge faster than any other method.
B) Offers a balance between the efficiency of batch GD and the robustness of SGD.
C) Does not require a loss function.
D) Is only applicable to linear models.
  • 2. "Transfer Learning" in deep learning involves:
A) Taking a model pre-trained on a large dataset (e.g., ImageNet) and fine-tuning it for a new, specific task with a smaller dataset.
B) Training a model from scratch on every new problem.
C) Using only unsupervised learning techniques.
D) Forgetting everything a model has learned.
  • 3. An "Autoencoder" is a type of neural network primarily used for:
A) Predicting continuous values in a regression task.
B) Reinforcement learning.
C) Unsupervised learning tasks like dimensionality reduction and data denoising.b
D) Supervised classification of images.
  • 4. The architecture of a typical autoencoder consists of:
A) An encoder that compresses the input and a decoder that reconstructs the input from the compression.
B) A single output neuron with a linear activation.
C) Only a single layer of perceptrons.
D) A convolutional layer followed by an RNN layer.
  • 5. In the context of model evaluation for classification, "Accuracy" is defined as:
A) The harmonic mean of precision and recall.
B) The proportion of positive identifications that were actually correct.
C) The proportion of total predictions that were correct.
D) The proportion of actual positives that were identified correctly.
  • 6. "Precision" is an important metric when:
A) You need a single metric that combines precision and recall.
B) You are evaluating a regression model.
C) The cost of false positives is high (e.g., in spam detection, where you don't want to flag legitimate emails as spam).
D) The cost of false negatives is high (e.g., in disease screening, where you don't want to miss a sick patient).
  • 7. "Recall" is an important metric when:
A) The cost of false positives is high (e.g., in spam detection).
B) You need a single metric that combines precision and recall.
C) The cost of false negatives is high (e.g., in disease screening, where you don't want to miss a sick patient).
D) You are evaluating a clustering model.
  • 8. The "F1 Score" is:
A) The difference between precision and recall.
B) A metric used exclusively for regression.
C) The arithmetic mean of precision and recall.
D) The harmonic mean of precision and recall, providing a single score that balances both concerns.
  • 9. For a regression model, the "Mean Squared Error" (MSE) measures:
A) The variance of the input features.
B) The accuracy of a classification model.
C) The average of the squares of the errors between predicted and actual values.
D) The total number of misclassified instances.
  • 10. The "ROC Curve" is a tool used to evaluate:
A) The loss of a regression model over time.
B) The architecture of a neural network.
C) The performance of a binary classification model at various classification thresholds.
D) The clustering quality of a K-means algorithm.
  • 11. "Area Under the ROC Curve" (AUC) provides an aggregate measure of performance across all possible classification thresholds. A perfect model has an AUC of:
A) 0.0.
B) 0.5.
C) -1.0.
D) 1.0.
  • 12. "K-fold Cross-Validation" is a technique used to:
A) Visualize high-dimensional data.
B) Increase the size of the training dataset.
C) Obtain a more robust estimate of model performance by training and evaluating the model K times on different splits of the data.
D) Replace the need for a separate test set.
  • 13. In the K-Nearest Neighbors (K-NN) algorithm for classification, the class of a new data point is determined by:
A) The majority vote among its K closest neighbors in the feature space.
B) A random selection from the training set.
C) A single, pre-defined rule.
D) The output of a linear function.u
  • 14. The parameter 'K' in the K-NN algorithm:
A) Is the number of features in the dataset.
B) Controls the model's flexibility. A small K can lead to overfitting, while a large K can lead to underfitting.
C) Is the learning rate for the algorithm.
D) Is always set to 1 for the best performance.
  • 15. "Principal Component Analysis" (PCA) works by:
A) Clustering data into K groups.
B) Finding new, uncorrelated dimensions (principal components) that capture the maximum variance in the data.u
C) Classifying data using a decision boundary.
D) Predicting a target variable using linear combinations of features.
  • 16. The first principal component in PCA is the direction in the feature space that:
A) Is randomly oriented.
B) Captures the least possible variance in the data.
C) Is perpendicular to all other components.
D) Captures the greatest possible variance in the data.
  • 17. "K-Means Clustering" aims to partition data into K clusters such that:
A) The between-cluster variance is minimized.
B) The within-cluster variance is minimized.
C) The data is projected onto a single dimension.
D) The data is perfectly classified into known labels.
  • 18. The "Elbow Method" is a heuristic used in K-Means to:
A) Determine the learning rate for gradient descent.
B) Initialize the cluster centroids.
C) Help choose the optimal number of clusters K by looking for a "bend" in the plot of within-cluster variance.
D) Evaluate the accuracy of a classification model.
  • 19. "Naive Bayes" classifiers are called "naive" because they:
A) Do not use probability in their predictions.
B) Always have the lowest possible accuracy.
C) Make a strong (naive) assumption that all features are conditionally independent given the class label.
D) Are very simple and cannot handle complex data.
  • 20. "Logistic Regression" is fundamentally a:
A) Dimensionality reduction technique.
B) Clustering algorithm for grouping unlabeled data.
C) Regression algorithm for predicting continuous values.
D) Classification algorithm that models the probability of a binary outcome using a logistic function.
  • 21. The output of a logistic regression model is a value between 0 and 1, which represents the:
A) Probability that the input belongs to a particular class.
B) Exact value of the target variable.
C) Distance to the decision boundary.
D) Number of features in the input.
  • 22. A "Random Forest" is an ensemble method that combines multiple:
A) Support Vector Machines.
B) Linear Regression models.
C) K-NN models.
D) Decision Trees to reduce overfitting and improve generalization.
  • 23. The "bagging" technique in a Random Forest helps to:
A) Reduce bias by making trees more complex.
B) Reduce variance by training individual trees on random subsets of the data and averaging their results
C) Increase the speed of a single decision tree.
D) Perform feature extraction like PCA.
  • 24. "Gradient Boosting" machines (e.g., XGBoost) are ensemble methods that:
A) Build models sequentially, where each new model corrects the errors of the previous ones.
B) Are exclusively used for unsupervised learning.
C) Do not require any parameter tuning.
D) Build all models independently and average them.
  • 25. The term "feature engineering" refers to:
A) The process of deleting all features from a dataset.
B) The automatic learning of features by a deep neural network.
C) The evaluation of a model's final performance.
D) The process of using domain knowledge to create new input features that make machine learning algorithms work better.
  • 26. "One-hot encoding" is a preprocessing technique used to:
A) Cluster similar data points together.
B) Convert categorical variables into a binary (0/1) format that can be provided to ML algorithms.
C) Normalize continuous numerical features.
D) Reduce the dimensionality of image data.
  • 27. "Feature scaling" (e.g., normalization or standardization) is often crucial for algorithms that:
A) Are used for clustering only.
B) Are based on tree-based models like Decision Trees and Random Forests.
C) Are used for association rule learning.
D) Are based on distance calculations or gradient descent, such as SVM and Neural Networks.
  • 28. The "curse of dimensionality" refers to the problem that:
A) Dimensionality reduction always improves model performance.
B) All datasets should have as many features as possible.
C) There are never enough features to train a good model.
D) As the number of features grows, the data becomes increasingly sparse, making it harder to find meaningful patterns.
  • 29. "Regularization" is a technique used to:
A) Speed up the training time of a model.
B) Prevent overfitting by adding a penalty term to the loss function that discourages complex models.
C) Increase the variance of a model.
D) Make models more complex to fit the training data better.
  • 30. L1 Regularization (Lasso) can often lead to:
A) Sparse models where the weights of less important features are driven to zero, effectively performing feature selection.
B) All features having non-zero weights.
C) Increased model complexity.
D) A decrease in model interpretability.
  • 31. "Hyperparameters" are:
A) The parameters that the model learns during training (e.g., weights in a neural network).
B) The output predictions of the model.
C) Configuration settings for the learning algorithm that are not learned from the data and must be set prior to training (e.g., learning rate, K in K-NN).
D) The input features of the model.
  • 32. The process of "Hyperparameter Tuning" involves:
A) Searching for the best combination of hyperparameters that results in the best model performance.
B) Training the model's internal weights.
C) Deploying the final model.
D) Cleaning the raw data.
  • 33. "Grid Search" is a common method for hyperparameter tuning that involves:
A) Randomly sampling hyperparameter combinations from a distribution.
B) Exhaustively searching over a specified set of hyperparameter values.
C) Using a separate neural network to predict the best hyperparameters.
D) Ignoring hyperparameters altogether.
  • 34. "Early Stopping" is a form of regularization that works by:
A) Stopping the training after a fixed, very short number of epochs.
B) Using a very small learning rate.
C) Starting the training process later than scheduled.
D) Halting the training process when performance on a validation set starts to degrade, indicating the onset of overfitting.
  • 35. A "Vanilla" neural network, also known as a Multilayer Perceptron (MLP), is typically composed of:
A) A single layer of neurons.
B) Convolutional layers for processing images.
C) Fully connected layers, where each neuron in one layer is connected to every neuron in the next layer.
D) Recurrent layers for processing sequences.
  • 36. The "softmax" activation function is commonly used in the output layer of a neural network for:
A) Multi-class classification problems, as it outputs a probability distribution over the possible classes.
B) Binary classification problems.
C) Regression problems.
D) Unsupervised learning problems.
  • 37. The "Adam" optimizer is an adaptive learning rate algorithm that is often preferred because it:
A) Is only used for unsupervised learning.
B) Combines the advantages of two other extensions of stochastic gradient descent, AdaGrad and RMSProp.
C) Is guaranteed to find the global minimum for any function.
D) Does not require any hyperparameters.
  • 38. "Batch Normalization" is a technique used to:
A) Replace the need for an activation function.
B) Improve the stability and speed of neural network training by normalizing the inputs to each layer.
C) Normalize the entire dataset before feeding it into the network.
D) Increase the batch size during training.
  • 39. The "confusion matrix" is a table that is used to describe the performance of a:
A) Regression model's accuracy.
B) Classification model on a set of test data for which the true values are known.
C) Clustering algorithm's group assignments.
D) Dimensionality reduction technique's effectiveness.
  • 40. In a confusion matrix, the "true positives" are the cases where:
A) The model incorrectly predicted the negative class.
B) The model correctly predicted the positive class.
C) The model incorrectly predicted the positive class.
D) The model correctly predicted the negative class.
  • 41. The problem of "imbalanced classes" occurs when:
A) The features are not scaled properly.
B) One class in the training data has significantly more examples than another, which can bias the model.
C) The model is too complex for the data.
D) The learning rate is set too high.
  • 42. A technique to address imbalanced classes is "SMOTE," which:
A) Ignores the minority class completely.
B) Combines all classes into one.
C) Generates synthetic examples for the minority class to balance the dataset.
D) Deletes examples from the majority class at random.
  • 43. "Reinforcement Learning" differs from supervised and unsupervised learning in that:
A) It is a simpler and less powerful approach.
B) It is only used for clustering unlabeled data.
C) It learns by interacting with an environment and receiving rewards or penalties for actions, without a labeled dataset.
D) It requires a fully labeled dataset for training.
  • 44. "Q-Learning" is a popular algorithm in reinforcement learning that learns:
A) A decision tree for classification.
B) A policy that tells an agent what action to take under what circumstances by learning a value function.
C) A clustering of possible actions.
D) The principal components of a state space.
  • 45. "Natural Language Processing" (NLP) often uses supervised learning for tasks like:
A) Sentiment analysis, where text is classified as positive, negative, or neutral.
B) Reducing the dimensionality of word vectors.
C) Generating new, original text without any input.
D) Grouping similar news articles without labels.
  • 46. "Word Embeddings" (like Word2Vec) are techniques that:
A) Are a type of clustering algorithm.
B) Represent words as dense vectors in a continuous space, capturing semantic meaning.
C) Represent words as simple one-hot encoded vectors.
D) Are used only for image classification.
  • 47. A "Generative Adversarial Network" (GAN) consists of two networks:
A) An Encoder and a Decoder for compression.
B) Two identical Convolutional Neural Networks.
C) A single, large Regression network.
D) A Generator and a Discriminator, which are trained in opposition to each other.
  • 48. The "Generator" in a GAN is responsible for:
A) Creating new, synthetic data that is indistinguishable from real data.
B) Classifying input images into categories.
C) Reducing the dimensionality of the input.
D) Discriminating between real and fake data.
  • 49. The "Discriminator" in a GAN is essentially a:
A) Regression model predicting a continuous value.
B) Binary classifier that tries to correctly label data as real (from the dataset) or fake (from the generator).
C) Clustering algorithm grouping similar images.
D) Dimensionality reduction technique.
Created with That Quiz — a math test site for students of all grade levels.