Databricks-Certified-Professional-Data-Scientistトレーニング最新認定問題をゲットDatabricks Certification合格目指せ2022年03月19日
認定トレーニングDatabricks-Certified-Professional-Data-Scientist試験問題集テストエンジン
質問 47
In which of the following scenario we can use naTve Bayes theorem for classification
- A. To identify whether a fruit is an orange or not based on features like diameter, color and shape
- B. Classify whether a given person is a male or a female based on the measured features. The features include height, weight and foot size.
- C. To classify whether an email is spam or not spam
正解: A,B,C
解説:
Explanation
naive Bayes classifiers have worked quite well in many real-world situations, famously document classification and spam filtering. They requires a small amount of training data to estimate the necessary parameters
質問 48
What is the best way to evaluate the quality of the model found by an unsupervised algorithm like k-means clustering, given metrics for the cost of the clustering (how well it fits the data) and its stability (how similar the clusters are across multiple runs over the same data)?
- A. The most stable clustering
- B. The lowest cost clustering
- C. The most stable clustering subject to a minimal cost constraint
- D. The lowest cost clustering subject to a stability constraint
正解: D
解説:
Explanation
There is a tradeoff between cost and stability in unsupervised learning. The more tightly you fit the data, the less stable the model will be, and vice versa. The idea is to find a good balance with more weight given to the cost. Typically a good approach is to set a stability threshold and select the model that achieves the lowest cost above the stability threshold.
質問 49
You are working with the Clustering solution of the customer datasets. There are almost 40 variables are available for each customer and almost 1.00,0000 customer's data is available. You want to reduce the number of variables for clustering, what would you do?
- A. You cannot discard any variable for creating clusters.
- B. You will randomly reduce the number of variables
- C. You can combine several variables in one variable
- D. You will find the correlation among the variables and from the highly co-related variables, you will be considering only one or two variables from it.
- E. You will find the correlation among the variables and from their variables are not co-related will be discarded.
正解: C,D
解説:
Explanation
When you are applying clustering technique and you find that there are quite a huge number of variables are available. Then it is better the find the co-relation among the variables and consider only one or two variables from the highly co-related variables. Because highly co-related variable will have the same effect, while creating the cluster. We can use scatter plot matrix among the variables to find the co-relation.
You can also combine several variables into a single variable. For example if you have two values in the dataset like Asset and Debt than by combining these two values like Debt to Asset ratio and use it while creating the cluster.
質問 50
You are having 1000 patients' data with the height and age. Where age in years and height in meters. You wanted to create cluster using this two attributes. You wanted to have near equal effect for both the age and height while creating the cluster. What you can do?
- A. You will be dividing both age and height with their respective standard deviation
- B. You will be converting each height value to centimeters
- C. You will be adding height with the numeric value 100
- D. You will be taking square root of height
正解: A,B
解説:
Explanation
When you see the data age in years would have values like 50, 60r 70 90 years etc. And while calculating distance from centroid maximum possible value can be 90-0 and its square will be 8100.
While using heights in meter can be 2-0.5(1.5) meters and its square will be 2.25 only. So you can see age has more effect than height. Hence bringing the height on same level you can convert it into centimeters. Can bring data upto 200 centimeters and then it be more effective like square of 200 maximum.
However there is another approach is to divide the each value with its standard deviation, which will not have impact of the units e.g. age/sd of the age, which results in value without unit. This can also help in reducing the effect of units.
質問 51
You have used k-means clustering to classify behavior of 100, 000 customers for a retail store. You decide to use household income, age, gender and yearly purchase amount as measures. You have chosen to use 8 clusters and notice that 2 clusters only have 3 customers assigned. What should you do?
- A. Decrease the number of measures used
- B. Identify additional measures to add to the analysis
- C. Increase the number of clusters
- D. Decrease the number of clusters
正解: D
解説:
Explanation
kmeans uses an iterative algorithm that minimizes the sum of distances from each object to its cluster centroid, over all clusters. This algorithm moves objects between clusters until the sum cannot be decreased further. The result is a set of clusters that are as compact and well-separated as possible. You can control the details of the minimization using several optional input parameters to kmeans, including ones for the initial values of the cluster centroids, and for the maximum number of iterations.
Clustering is primarily an exploratory technique to discover hidden structures of the data: possibly as a prelude to more focused analysis or decision processes. Some specific applications of k-means are image processing^ medical and customer segmentation. Clustering is often used as a lead-in to classification. Once the clusters are identified, labels can be applied to each cluster to classify each group based on its characteristics. Marketing and sales groups use k-means to better identify customers who have similar behaviors and spending patterns.
質問 52
Classification and regression are examples of___________.
- A. Density estimation
- B. Clustering
- C. supervised learning
- D. un-supervised learning
正解: C
解説:
Explanation
In classification, our job is to predict what class an instance of data should fall into. Another task in machine learning is regression. Regression is the prediction of a numeric value. Most people have probably seen an example of regression with a best-fit line drawn through some data points to generalize the data points.
Classification and regression are examples of supervised learning. This set of problems is known as supervised because we're telling the algorithm what to predict.
質問 53
Question-34. Stories appear in the front page of Digg as they are "voted up" (rated positively) by the community. As the community becomes larger and more diverse, the promoted stories can better reflect the average interest of the community members. Which of the following technique is used to make such recommendation engine?
- A. Naive Bayes classifier
- B. Collaborative filtering
- C. Logistic Regression
- D. Content-based filtering
正解: B
解説:
Explanation
One scenario of collaborative filtering application is to recommend interesting or popular information as judged by the community. As a typical example, stories appear in the front page of Digg as they are "voted up" (rated positively) by the community. As the community becomes larger and more diverse, the promoted stories can better reflect the average interest of the community members.
質問 54
A researcher is interested in how variables, such as GRE (Graduate Record Exam scores), GPA (grade point average) and prestige of the undergraduate institution, effect admission into graduate school. The response variable, admit/don't admit, is a binary variable.
Above is an example of
- A. Logistic Regression
- B. Linear Regression
- C. Maximum likelihood estimation
- D. Recommendation system
- E. Hierarchical linear models
正解: A
解説:
Explanation
Logistic regression
Pros: Computationally inexpensive, easy to implement, knowledge representation easy to interpret Cons: Prone to underfitting, may have low accuracy Works with: Numeric values, nominal values
質問 55
In which of the scenario you can use the regression to predict the values
- A. Samsung can use it for mobile sales forecast
- B. Probability of the celebrity divorce
- C. Mobile companies can use it to forecast manufacturing defects
- D. Only 1 and 2
- E. All 1 ,2 and 3
正解: E
解説:
Explanation
Regression is a tool which Companies may use this for things such as sales forecasts or forecasting manufacturing defects. Another creative example is predicting the probability of celebrity divorce.
質問 56
You have modeled the datasets with 5 independent variables called A,B,C,D and E having relationships which is not dependent each other, and also the variable A,B and C are continuous and variable D and E are discrete (mixed mode).
Now you have to compute the expected value of the variable let say A, then which of the following computation you will prefer
- A. Transformation
- B. Integration
- C. Differentiation
- D. Generalization
正解: B
解説:
Explanation
Text Description automatically generated
Text Description automatically generated
Text Description automatically generated
質問 57
You are working on a email spam filtering assignment, while working on this you find there is new word e.g.
HadoopExam comes in email, and in your solutions you never come across this word before, hence probability of this words is coming in either email could be zero. So which of the following algorithm can help you to avoid zero probability?
- A. All of the above
- B. Logistic Regression
- C. Laplace Smoothing
- D. Naive Bayes
正解: C
解説:
Explanation
Laplace smoothing is a technique for parameter estimation which accounts for unobserved events. It is more robust and will not fail completely when data that has never been observed in training shows up.
質問 58
Which method is used to solve for coefficients bO, b1, ... bn in your linear regression model:
- A. Integer programming
- B. Apriori Algorithm
- C. Ridge and Lasso
- D. Ordinary Least squares
正解: D
解説:
Explanation : RY = b0 + b1x1+b2x2+ .... +bnxn
In the linear model, the bi's represent the unknown p parameters. The estimates for these unknown parameters are chosen so that, on average, the model provides a reasonable estimate of a person's income based on age and education. In other words, the fitted model should minimize the overall error between the linear model and the actual observations. Ordinary Least Squares (OLS) is a common technique to estimate the parameters
質問 59
Which technique you would be using to solve the below problem statement? "What is the probability that individual customer will not repay the loan amount?"
- A. Clustering
- B. Hypothesis testing
- C. Logistic Regression
- D. Linear Regression
- E. Classification
正解: C
質問 60
Refer to Exhibit
In the exhibit, the x-axis represents the derived probability of a borrower defaulting on a loan. Also in the exhibit, the pink represents borrowers that are known to have not defaulted on their loan, and the blue represents borrowers that are known to have defaulted on their loan. Which analytical method could produce the probabilities needed to build this exhibit?
- A. Discriminant Analysis
- B. Logistic Regression
- C. Association Rules
- D. Linear Regression
正解: B
質問 61
You are building a classifier off of a very high-dimensiona data set similar to shown in the image with 5000 variables (lots of columns, not that many rows). It can handle both dense and sparse input. Which technique is most suitable, and why?
- A. Naive Bayes, because Bayesian methods act as regularlizers
- B. k-nearest neighbors, because it uses local neighborhoods to classify examples
- C. Random forest because it is an ensemble method
- D. Logistic regression with L1 regularization, to prevent overfitting
正解: D
解説:
Explanation
Logistic regression is widely used in machine learning for classification problems. It is well-known that regularization is required to avoid over-fitting, especially when there is a only small number of training examples, or when there are a large number of parameters to be learned. In particular L1 regularized logistic regression is often used for feature selection, and has been shown to have good generalization performance in the presence of many irrelevant features. (Ng 2004; Goodman 2004) Unregularized logistic regression is an unconstrained convex optimization problem with a continuously differentiate objective function. As a consequence, it can be solved fairly efficiently with standard convex optimization methods, such as Newton's method or conjugate gradient. However, adding the L1 regularization makes the optimization problem com-putationally more expensive to solve. If the L1 regulariza-tion is enforced by an L1 norm constraint on the parameLogistic regression is a classifier and L1 regularization tends to produce models that ignore dimensions of the input that are not predictive. This is particularly useful when the input contains many dimensions, k-nearest neighbors classification is also a classification technique, but relies on notions of distance. In a high-dimensional space, most every data point is "far" from others (the curse of dimensionality) and so these techniques break down. Naive Bayes is not inherently regularizing. Random forests represent an ensemble method; but an ensemble method is not necessarily more suitable to high-dimensional data.
Practically, I think the biggest reasons for regularization are 1) to avoid overfitting by not generating high coefficients for predictors that are sparse. 2) to stabilize the estimates especially when there's collinearity in the data.
1) is inherent in the regularization framework. Since there are two forces pulling each other in the objective function, if there's no meaningful loss reduction, the increased penalty from the regularization term wouldn't improve the overall objective function. This is a great property since a lot of noise would be automatically filtered out from the model. To give you an example for 2), if you have two predictors that have same values, if you just run a regression algorithm on it since the data matrix is singular your beta coefficients will be Inf if you try to do a straight matrix inversion. But if you add a very small regularization lambda to it, you will get stable beta coefficients with the coefficient values evenly divided between the equivalent two variables. For the difference between L1 and L2, the following graph demonstrates why people bother to have L1 since L2 has such an elegant analytical solution and is so computationally straightforward. Regularized regression can also be represented as a constrained regression problem (since they are Lagrangian equivalent). The implication of this is that the L1 regularization gives you sparse estimates. Namely, in a high dimensional space, you got mostly zeros and a small number of non-zero coefficients. This is huge since it incorporates variable selection to the modeling problem. In addition, if you have to score a large sample with your model, you can have a lot of computational savings since you don't have to compute features(predictors) whose coefficient is 0. I personally think L1 regularization is one of the most beautiful things in machine learning and convex optimization. It is indeed widely used in bioinformatics and large scale machine learning for companies like Facebook, Yahoo, Google and Microsoft.
質問 62
Which of the following is a correct example of the target variable in regression (supervised learning)?
- A. Nominal values like true, false
- B. Reptile, fish, mammal, amphibian, plant, fungi
- C. All of the above
- D. Infinite number of numeric values, such as 0.100, 42.001, 1000.743..
正解: C
解説:
Explanation
We address two cases of the target variable. The first case occurs when the target variable can take only nominal values: true or false; reptile, fish: mammal, amphibian, plant, fungi. The second case of classification occurs when the target variable can take an infinite number of numeric values, such as 0.100, 42.001,
1000.743, .... This case is called regression.
質問 63
You have collected the 100's of parameters about the 1000's of websites e.g. daily hits, average time on the websites, number of unique visitors, number of returning visitors etc. Now you have find the most important parameters which can best describe a website, so which of the following technique you will use
- A. Clustering
- B. PCA (Principal component analysis)
- C. Logistic Regression
- D. Linear Regression
正解: B
解説:
Explanation
Principal component analysis . or PCA, is a technique for taking a dataset that is in the form of a set of tuples representing points in a high-dimensional space and finding the dimensions along which the tuples line up best. The idea is to treat the set of tuples as a matrix M and find the eigenvectors for MMT or M T M . The matrix of these eigenvectors can be thought of as a rigid rotation in a high-dimensional space. When you apply this transformation to the original data, the axis corresponding to the principal eigenvector is the one along which the points are most "spread out,11 More precisely this axis is the one along which the variance of the data is maximized. Put another way, the points can best be viewed as lying along this axis, with small deviations from this axis.
質問 64
Which of the following skills a data scientists required?
- A. He should be creative
- B. Should be very good at mathematics and statistic
- C. He should possess database administrative skills.
- D. Should possess good programming skills
- E. Web designing to represent best visuals of its results from algorithm.
正解: A,B,D
解説:
Explanation
Yes a data scientists should have combination of skills like to solve the complex problem he should be creative as well as able to find new solutions and use of existing data. And solve the problem skills required are programming as currently we see SAS, R: Python, Spark, Java and SPSS even day by day new technologies are coming.
To apply various existing and new algorithm using Machine Learning, or Al it require good mathematics and statistics skills (Where the programmer feels, weaknesses). Another skill required is using visualization techniques like Qlik, Tableau etc
質問 65
You are analyzing data in order to build a classifier model. You discover non-linear data and discontinuities that will affect the model. Which analytical method would you recommend?
- A. Decision Trees
- B. Logistic Regression
- C. Linear Regression
- D. ARIMA
正解: A
解説:
Explanation
A decision tree is a flowchart-like structure in which each internal node represents a "test" on an attribute (e.g.
whether a coin flip comes up heads or tails), each branch represents the outcome of the test and each leaf node represents a class label (decision taken after computing all attributes). The paths from root to leaf represents classification rules.
In decision analysis a decision tree and the closely related influence diagram are used as a visual and analytical decision support tool, where the expected values (or expected utility) of competing alternatives are calculated.
A decision tree consists of 3 types of nodes:
1. Decision nodes - commonly represented by squares
2. Chance nodes - represented by circles
3. End nodes - represented by triangles
Decision trees are commonly used in operations research, specifically in decision analysis, to help identify a strategy most likely to reach a goal. If in practice decisions have to be taken online with no recall under incomplete knowledge, a decision tree should be paralleled by a probability model as a best choice model or online selection model algorithm. Another use of decision trees is as a descriptive means for calculating conditional probabilities.
Decision trees, influence diagrams, utility functions, and other decision analysis tools and methods are taught to undergraduate students in schools of business, health economics, and public health, and are examples of operations research or management science methods.
質問 66
You are using k-means clustering to classify heart patients for a hospital. You have chosen Patient Sex, Height, Weight, Age and Income as measures and have used 3 clusters. When you create a pair-wise plot of the clusters, you notice that there is significant overlap between the clusters. What should you do?
- A. Remove one of the measures
- B. Identify additional measures to add to the analysis
- C. Increase the number of clusters
- D. Decrease the number of clusters
正解: D
質問 67
Of all the smokers in a particular district, 40% prefer brand A and 60% prefer brand B.Of those smokers who prefer brand A. 30% are females, and of those who prefer brand B.40% are female. What is the probability that a randomly selected smoker prefers brand A, given that the person selected is a female?
Which of the following is a best way to solve this problem?
- A. None of the above
- B. Binomial Distribution
- C. Bays Theorem
- D. Poisson Distribution
正解: C
質問 68
Which of the following steps you will be using in the discovery phase?
- A. What all tools are required, in the project?
- B. What Unix server capacity required?
- C. Analyze the Raw data and its format and structure.
- D. What all are the data sources for the project?
- E. What is the network capacity required
正解: A,B,C,D,E
解説:
Explanation
During the discovery phase you need to find how much resources are required as early as possible and for that even you can involve various stakeholders like Software engineering team, DBAs, Network engineers, System administrators etc. for your requirement and these resources are already available or you need to procure them. Also, what would be source of the data? What all tools and software's are required to execute the same?
質問 69
Support vector machines (SVMs) are a set of supervised learning methods used for
- A. Linear classification
- B. Regression
- C. Non-linear classification
正解: A,B,C
解説:
Explanation
In machine learning, support vector machines (SVMs). also support vector networks[1]) are supervised learning models with associated learning algorithms that analyze data and recognize patterns^ used for classification and regression analysis. In addition to performing linear classification, SVMs can efficiently perform a non-linear classification using what is called the kernel tricky implicitly mapping their inputs into high-dimensional feature spaces.
質問 70
Projecting a multi-dimensional dataset onto which vector has the greatest variance?
- A. second principal component
- B. not enough information given to answer
- C. second eigenvector
- D. first eigenvector
- E. first principal component
正解: E
解説:
Explanation
The method based on principal component analysis (PCA) evaluates the features according to the projection of the largest eigenvector of the correlation matrix on the initial dimensions, the method based on Fisher's linear discriminant analysis evaluates. Them according to the magnitude of the components of the discriminant vector.
The first principal component corresponds to the greatest variance in the data, by definition. If we project the data onto the first principal component line, the data is more spread out (higher variance) than if projected onto any other line, including other principal components.
質問 71
......
Databricks Databricks-Certified-Professional-Data-Scientist 認定試験の出題範囲:
| トピック | 出題範囲 |
|---|---|
| トピック 1 |
|
| トピック 2 |
|
| トピック 3 |
|
| トピック 4 |
|
| トピック 5 |
|
合格確定、ガイドで準備Databricks-Certified-Professional-Data-Scientist試験:https://www.goshiken.com/Databricks/Databricks-Certified-Professional-Data-Scientist-mondaishu.html