試験準備には欠かさない!Databricks-Certified-Professional-Data-Scientist問題解答でDatabricks-Certified-Professional-Data-Scientist試験問題集
リアルDatabricks Databricks-Certified-Professional-Data-Scientist試験問題 [更新されたのは2022年]
質問 84
Which of the following is a Continuous Probability Distributions?
- A. Negative binomial distribution
- B. Normal probability distribution
- C. Binomial probability distribution
- D. Poisson probability distribution
正解: B
質問 85
You have collected the 100's of parameters about the 1000's of websites e.g. daily hits, average time on the websites, number of unique visitors, number of returning visitors etc. Now you have find the most important parameters which can best describe a website, so which of the following technique you will use
- A. Clustering
- B. PCA (Principal component analysis)
- C. Logistic Regression
- D. Linear Regression
正解: B
解説:
Explanation
Principal component analysis . or PCA, is a technique for taking a dataset that is in the form of a set of tuples representing points in a high-dimensional space and finding the dimensions along which the tuples line up best. The idea is to treat the set of tuples as a matrix M and find the eigenvectors for MMT or M T M . The matrix of these eigenvectors can be thought of as a rigid rotation in a high-dimensional space. When you apply this transformation to the original data, the axis corresponding to the principal eigenvector is the one along which the points are most "spread out,11 More precisely this axis is the one along which the variance of the data is maximized. Put another way, the points can best be viewed as lying along this axis, with small deviations from this axis.
質問 86
You are designing a recommendation engine for a website where the ability to generate more personalized recommendations by analyzing information from the past activity of a specific user, or the history of other users deemed to be of similar taste to a given user. These resources are used as user profiling and helps the site recommend content on a user-by-user basis. The more a given user makes use of the system, the better the recommendations become, as the system gains data to improve its model of that user. What kind of this recommendation engine is ?
- A. Naive Bayes classifier
- B. Collaborative filtering
- C. Logistic Regression
- D. Content-based filtering
正解: B
解説:
Explanation
Another aspect of collaborative filtering systems is the ability to generate more personalized recommendations by analyzing information from the past activity of a specific user, or the history of other users deemed to be of similar taste to a given user. These resources are used as user profiling and help the site recommend content on a user-by-user basis. The more a given user makes use of the system, the better the recommendations become, as the system gains data to improve its model of that user
質問 87
What is one modeling or descriptive statistical function in MADlib that is typically not provided in a standard relational database?
- A. Linear regression
- B. Expected value
- C. Quantiles
- D. Variance
正解: A
解説:
Explanation
Linear regression models a linear relationship of a scalar dependent variable y to one or more explanatory independent variables x to build a model of coefficients.
質問 88
You are working in a classification model for a book, written by HadoopExam Learning Resources and decided to use building a text classification model for determining whether this book is for Hadoop or Cloud computing. You have to select the proper features (feature selection) hence, to cut down on the size of the feature space, you will use the mutual information of each word with the label of hadoop or cloud to select the 1000 best features to use as input to a Naive Bayes model. When you compare the performance of a model built with the 250 best features to a model built with the 1000 best features, you notice that the model with only 250 features performs slightly better on our test data.
What would help you choose better features for your model?
- A. Evaluate a model that only includes the top 100 words
- B. Decrease the size of our training data
- C. Include least mutual information with other selected features as a feature selection criterion
- D. Include the number of times each of the words appears in the book in your model
正解: C
解説:
Explanation
Correlation measures the linear relationship (Pearson's correlation) or monotonic relationship (Spearman's correlation) between two variables, X and Y.
Mutual information is more general and measures the reduction of uncertainty in Y after observing X.
It is the KL distance between the joint density and the product of the individual densities. So Ml can measure non-monotonic relationships and other more complicated relationships Mutual information is a quantification of the dependency between random variables. It is sometimes contrasted with linear correlation since mutual information captures nonlinear dependence.
Features with high mutual information with the predicted value are good. However a feature may have high mutual information because it is highly correlated with another feature that has already been selected.
Choosing another feature with somewhat less mutual information with the predicted value, but low mutual information with other selected features, may be more beneficial. Hence it may help to also prefer features that are less redundant with other selected features.
質問 89
Select the correct objectives of principal component analysis
- A. To discover the dimensionality of the data set
- B. To reduce the dimensionality of the data set
- C. All 1, 2 and 3
- D. To identify new meaningful underlying variables
- E. Only 1 and 2
正解: C
解説:
Explanation
Principal component analysis (PCA) involves a mathematical procedure that transforms a number of (possibly) correlated variables into a (smaller) number of uncorrelated variables called principal components. The first principal component accounts for as much of the variability in the data as possible: and each succeeding component accounts for as much of the remaining variability as possible.
Objectives of principal component analysis
1. To discover or to reduce the dimensionality of the data set.
2. To identify new meaningful underlying variables.
質問 90
Find out the classifier which assumes independence among all its features?
- A. Neural networks
- B. Random forests
- C. Naive Bayes
- D. Linear Regression
正解: C
解説:
Explanation
A Bayes classifier is a simple probabilistic classifier based on applying Bayes' theorem (from Bayesian statistics) with strong (naive) independence assumptions. A more descriptive term for the underlying probability model would be "independent feature model".
A Bayes classifier is a simple probabilistic classifier based on applying Bayes' theorem (from Bayesian statistics) with strong (naive) independence assumptions. A more descriptive term for the underlying probability model would be "independent feature model".
In simple terms, a naive Bayes classifier assumes that the presence (or absence) of a particular feature of a class is unrelated to the presence (or absence) of any other feature. For example, a fruit may be considered to be an apple if it is red, round, and about 4" in diameter Even if these features depend on each other or upon the existence of the other features, a naive Bayes classifier considers all of these properties to independently contribute to the probability that this fruit is an apple.
質問 91
What are the key outcomes of the successful analytical projects?
- A. Presentation for Project Sponsors
- B. Code of the model
- C. Presentations for the Analysts
- D. Technical specifications
正解: A,B,C,D
解説:
Explanation
When your analytical project successfully completed they come up with the following at the end of the projects. Presentations- You will be having presentations like for the all the stakeholders, generally these presentation will help seniors executives to make better decisions. Similarly you would be creating presentations for the other teams like analysts various visuals you would be creating like ROC Curves, Heat Maps, and Bar Charts etc.
Whatever tools you have used like SAS, R, or Python then accordingly code was developed and you will get that code as one of the outcome. Also you would have created a technical specifications for implementing the codes.
質問 92
You are analyzing data in order to build a classifier model. You discover non-linear data and discontinuities that will affect the model. Which analytical method would you recommend?
- A. Decision Trees
- B. Logistic Regression
- C. Linear Regression
- D. ARIMA
正解: A
解説:
Explanation
A decision tree is a flowchart-like structure in which each internal node represents a "test" on an attribute (e.g.
whether a coin flip comes up heads or tails), each branch represents the outcome of the test and each leaf node represents a class label (decision taken after computing all attributes). The paths from root to leaf represents classification rules.
In decision analysis a decision tree and the closely related influence diagram are used as a visual and analytical decision support tool, where the expected values (or expected utility) of competing alternatives are calculated.
A decision tree consists of 3 types of nodes:
1. Decision nodes - commonly represented by squares
2. Chance nodes - represented by circles
3. End nodes - represented by triangles
Decision trees are commonly used in operations research, specifically in decision analysis, to help identify a strategy most likely to reach a goal. If in practice decisions have to be taken online with no recall under incomplete knowledge, a decision tree should be paralleled by a probability model as a best choice model or online selection model algorithm. Another use of decision trees is as a descriptive means for calculating conditional probabilities.
Decision trees, influence diagrams, utility functions, and other decision analysis tools and methods are taught to undergraduate students in schools of business, health economics, and public health, and are examples of operations research or management science methods.
質問 93
Reducing the data from many features to a small number so that we can properly visualize it in two or three dimensions. It is done in_______
- A. un-supervised learning
- B. Support vector machines
- C. supervised learning
- D. k-Nearest Neighbors
正解: A
解説:
Explanation
The opposite of supervised learning is a set of tasks known as unsupervised learning. In unsupervised learning, there's no label or target value given for the data. A task where we group similar items together is known as clustering. In unsupervised learning, we may also want to find statistical values that describe the data. This is known as density estimation. Another task of unsupervised learning may be reducing the data from many features to a small number so that we can properly visualize it in two or three dimensions
質問 94
Logistic regression is a model used for prediction of the probability of occurrence of an event. It makes use of several variables that may be......
- A. Both 1 and 2 are correct
- B. Categorical
- C. Numerical
- D. None of the 1 and 2 are correct
正解: A
解説:
Explanation
Logistic regression is a model used for prediction of the probability of occurrence of an event. It makes use of several predictor variables that may be either numerical or categories.
質問 95
Select the statement which applies correctly to the Naive Bayes
- A. Sensitive to how the input data is prepared
- B. Works with a small amount of data
- C. Works with nominal values
正解: A,B,C
質問 96 
The figure below shows a plot of the data of a data matrix M that is 1000 x 2. Which line represents the first principal component?
- A. Neither
- B. yellow
- C. blue
正解: C
解説:
Explanation
Principal component analysis (PCA) involves a mathematical procedure that transforms a number of (possibly) correlated variables into a (smaller) number of uncorrelated variables called principal components. The first principal component accounts for as much of the variability in the data as possible, and each succeeding component accounts for as much of the remaining variability as possible.
The first principal component corresponds to the greatest variance in the data. The blue line is evidently this first principal component, because if we project the data onto the blue line, the data is more spread out (higher variance) than if projected onto any other line, including the yellow one.
質問 97
Which of the following statement true with regards to Linear Regression Model?
- A. Ordinary Least Square is a sum of the squared individual distance between each point and the fitted line of regression model.
- B. In Linear model, it tries to find multiple lines which can approximate the relationship between the outcome and input variables.
- C. Ordinary Least Square can be used to estimates the parameters in linear model
- D. Ordinary Least Square is a sum of the individual distance between each point and the fitted line of regression model.
正解: A,C
解説:
Explanation
Linear regression model are represented using the below equation
Where B(0) is intercept and B(1) is a slope. As B(0) and B(1) changes then fitted line also shifts accordingly on the plot. The purpose of the Ordinary Least Square method is to estimates these parameters B(0) and B(1).
And similarly it is a sum of squared distance between the observed point and the fitted line. Ordinary least squares (OLS) regression minimizes the sum of the squared residuals. A model fits the data well if the differences between the observed values and the model's predicted values are small and unbiased.
質問 98
Your company has organized an online campaign for feedback on product quality and you have all the responses for the product reviews, in the response form people have check box as well as text field. Now you know that people who do not fill in or write non-dictionary word in the text field are not considered valid feedback. People who fill in text field with proper English words are considered valid response. Which of the following method you should not use to identify whether the response is valid or not?
- A. Random Decision Forests
- B. Logistic Regression
- C. Any one of the above
- D. Naive Bayes
正解: C
解説:
Explanation
In this problem you have been given high-dimensional independent variables like yeS; nO; no English words , test results etc. and you have to predict either valid or not valid (One of two). So all of the below technique can be applied to this problem.
* Support vector machines
* Naive Bayes
* Logistic regression
* Random decision forests
質問 99
Suppose there are three events then which formula must always be equal to P(E1|E2,E3)?
- A. P(E1,E2|E3)P(E3)
- B. P(E1,E2;E3)/P(E2,E3)
- C. P(E1,E2|E3)P(E2|E3)P(E3)
- D. P(E1,E2,E3)P(E2)P(E3)
- E. P(E1,E2,E3)P(E1)/P(E2:E3)
正解: B
解説:
Explanation
This is an application of conditional probability: P(E1,E2)=P(E1|E2)P(E2). so P(E1|E2) = P(E1.E2)/P(E2) P(E1,E2,E3)/P(E2,E3) If the events are A and B respectively, this is said to be "the probability of A given B" It is commonly denoted by P(A|B): or sometimes PB(A). In case that both "A" and "B" are categorical variables, conditional probability table is typically used to represent the conditional probability.
質問 100
Which of the following statement is true for the R square value in the regression model?
- A. R-squared never decreases upon adding more independent variables.
- B. When R square =0, all the residual are equal to 1
- C. R square can be increased by adding more variables to the model.
- D. When R square =1 , all the residuals are equal to 0
正解: A,C,D
質問 101
In which lifecycle stage are appropriate analytical techniques determined?
- A. Data preparation
- B. Discovery
- C. Model building
- D. Model planning
正解: D
解説:
Explanation
In Phase 3, the data science team identifies candidate models to apply to the data for clustering, classifying, or finding relationships in the data depending on the goal of the project, It is during this phase that the team refers to the hypotheses developed in Phase 1, when they first became acquainted with the data and understanding the business problems or domain area. These hypotheses help the team frame the analytics to execute in Phase
4 and select the right methods to achieve its objectives.
Some of the activities to consider in this phase include the following: Assess the structure of the datasets. The structure of the datasets is one factor that dictates the tools and analytical techniques for the next phase.
Depending on whether the team plans to analyze textual data or transactional data, for example, different tools and approaches are required.
Ensure that the analytical techniques enable the team to meet the business objectives and accept or reject the working hypotheses. Determine if the situation warrants a single model or a series of techniques as part of a larger analytic workflow. A few example models include association rules and logistic regression Other tools, such as Alpine Miner, enable users to set up a series of steps and analyses and can serve as a front-end user interface (Ul) for manipulating Big Data sources in PostgreSQL.
質問 102
You are creating a Classification process where input is the income, education and current debt of a customer, what could be the possible output of this process.
- A. Percentage of the customer loan repayment capability
- B. Percentage of the customer should be given loan or not
- C. Probability of the customer default on loan repayment
- D. The output might be a risk class, such as "good", "acceptable", "average", or "unacceptable".
正解: D
解説:
Explanation
Classification is the process of using several inputs to produce one or more outputs. For example the input might be the income, education and current debt of a customer The output might be a risk class, such as
"good", "acceptable", "average", or "unacceptable". Contrast this to regression where the output is a number not a class.
質問 103
You are working on a Data Science project and during the project you have been gibe a responsibility to interview all the stakeholders in the project. In which phase of the project you are?
- A. Operationnalise the models
- B. Creating Models
- C. Data Preparations
- D. Executing Models
- E. Creating visuals from the outcome
- F. Discovery
正解: F
解説:
Explanation
During the discovery phase you will be interviewing all the project stakeholders because they would be having quite a good amount of knowledge for the problem domain you will be working and you also interviewing project sponsors you will get to know what all are the expectations once project get completed. Hence, you will be noting down all the expectations from the project as well as you will be using their expertise in the domain.
質問 104
Your customer provided you with 2. 000 unlabeled records three groups. What is the correct analytical method to use?
- A. Logistic regression
- B. Linear regression
- C. Naive Bayesian classification
- D. K-means clustering
- E. Semi Linear Regression
正解: D
解説:
Explanation
k-means clustering is a method of vector quantization^ originally from signal processing, that is popular for cluster analysis in data mining, k-means clustering aims to partition n observations into k clusters in which each observation belongs to the cluster with the nearest mean, serving as a prototype of the cluster This results in a partitioning of the data space into Voronoi cells.
The problem is computationally difficult (NP-hard); however there are efficient heuristic algorithms that are commonly employed and converge quickly to a local optimum. These are usually similar to the expectation-maximization algorithm for mixtures of Gaussian distributions via an iterative refinement approach employed by both algorithms. Additionally they both use cluster centers to model the data; however k-means clustering tends to find clusters of comparable spatial extent, while the expectation-maximization mechanism allows clusters to have different shapes.
The algorithm has nothing to do with and should not be confused with k-nearest neighbor another popular machine learning technique.
質問 105
Spam filtering of the emails is an example of
- A. 2 and 3 are correct
- B. Clustering
- C. 1 and 3 are correct
- D. Unsupervised learning
- E. Supervised learning
正解: E
解説:
Explanation
Clustering is an example of unsupervised learning. The clustering algorithm finds groups within the data without being told what to look for upfront. This contrasts with classification, an example of supervised machine learning, which is the process of determining to which class an observation belongs. A common application of classification is spam filtering. With spam filtering we use labeled data to train the classifier:
e-mails marked as spam or ham.
質問 106
Feature Hashing approach is "SGD-based classifiers avoid the need to predetermine vector size by simply picking a reasonable size and shoehorning the training data into vectors of that size" now with large vectors or with multiple locations per feature in Feature hashing?
- A. Is a problem with accuracy
- B. It is easy to understand what classifier is doing
- C. Is a problem with accuracy as well as hard to understand what classifier us doing
- D. It is hard to understand what classifier is doing
正解: D
解説:
Explanation
FEATURE HASHING
SGD-based classifiers avoid the need to predetermine vector size by simply picking a reasonable size and shoehorning the training data into vectors of that size. This approach is known as feature hashing. The shoehorning is done by picking one or more locations by using a hash of the name of the variable for continuous variables or a hash of the variable name and the category name or word for categorical, text*like, or word-like data.
This hashed feature approach has the distinct advantage of requiring less memory and one less pass through the training data, but it can make it much harder to reverse engineer vectors to determine which original feature mapped to a vector location. This is because multiple features may hash to the same location. With large vectors or with multiple locations per feature, this isn't a problem for accuracy but it can make it hard to understand what a classifier is doing.
An additional benefit of feature hashing is that the unknown and unbounded vocabularies typical of word-like variables aren't a problem.
質問 107
......
Databricks Databricks-Certified-Professional-Data-Scientist 認定試験の出題範囲:
| トピック | 出題範囲 |
|---|---|
| トピック 1 |
|
| トピック 2 |
|
| トピック 3 |
|
Databricks-Certified-Professional-Data-Scientist合格させる試験問題集には更新されたのは2022年:https://www.goshiken.com/Databricks/Databricks-Certified-Professional-Data-Scientist-mondaishu.html