質問 1:A data scientist wants to use Spark ML to one-hot encode the categorical features in their PySpark DataFrame features_df. A list of the names of the string columns is assigned to the input_columns variable.
They have developed this code block to accomplish this task:

The code block is returning an error.
Which of the following adjustments does the data scientist need to make to accomplish this task?
A. They need to use Stringlndexer prior to one-hot encodinq the features.
B. They need to use VectorAssembler prior to one-hot encoding the features.
C. They need to remove the line with the fit operation.
D. They need to specify the method parameter to the OneHotEncoder.
正解:A
解説: (Topexam メンバーにのみ表示されます)
質問 2:A machine learning engineer is trying to scale a machine learning pipeline pipeline that contains multiple feature engineering stages and a modeling stage. As part of the cross-validation process, they are using the following code block:

A colleague suggests that the code block can be changed to speed up the tuning process by passing the model object to the estimator parameter and then placing the updated cv object as the final stage of the pipeline in place of the original model.
Which of the following is a negative consequence of the approach suggested by the colleague?
A. The cross-validation process will no longer be reproducible
B. The model will take longer to train for each unique combination of hvperparameter values
C. The feature engineering stages will be computed using validation data
D. The model will be refit one more per cross-validation fold
E. The cross-validation process will no longer be
正解:C
解説: (Topexam メンバーにのみ表示されます)
質問 3:An organization is developing a feature repository and is electing to one-hot encode all categorical feature variables. A data scientist suggests that the categorical feature variables should not be one-hot encoded within the feature repository.
Which of the following explanations justifies this suggestion?
A. One-hot encoding is not supported by most machine learning libraries.
B. One-hot encoding is a potentially problematic categorical variable strategy for some machine learning algorithms.
C. One-hot encoding is dependent on the target variable's values which differ for each application.
D. One-hot encoding is computationally intensive and should only be performed on small samples of training sets for individual machine learning problems.
E. One-hot encoding is not a common strategy for representing categorical feature variables numerically.
正解:B
解説: (Topexam メンバーにのみ表示されます)
質問 4:A machine learning engineer would like to develop a linear regression model with Spark ML to predict the price of a hotel room. They are using the Spark DataFrame train_df to train the model.
The Spark DataFrame train_df has the following schema:

The machine learning engineer shares the following code block:

Which of the following changes does the machine learning engineer need to make to complete the task?
A. They need to call the transform method on train df
B. They need to convert the features column to be a vector
C. They do not need to make any changes
D. They need to split the features column out into one column for each feature
E. They need to utilize a Pipeline to fit the model
正解:B
解説: (Topexam メンバーにのみ表示されます)
質問 5:A machine learning engineer wants to parallelize the inference of group-specific models using the Pandas Function API. They have developed the apply_model function that will look up and load the correct model for each group, and they want to apply it to each group of DataFrame df.
They have written the following incomplete code block:

Which piece of code can be used to fill in the above blank to complete the task?
A. groupedApplyInPandas
B. applyInPandas
C. predict
D. mapInPandas
正解:B
解説: (Topexam メンバーにのみ表示されます)
質問 6:A data scientist has developed a machine learning pipeline with a static input data set using Spark ML, but the pipeline is taking too long to process. They increase the number of workers in the cluster to get the pipeline to run more efficiently. They notice that the number of rows in the training set after reconfiguring the cluster is different from the number of rows in the training set prior to reconfiguring the cluster.
Which of the following approaches will guarantee a reproducible training and test set for each model?
A. Manually partition the input data
B. Write out the split data sets to persistent storage
C. Set a speed in the data splitting operation
D. Manually configure the cluster
正解:B
解説: (Topexam メンバーにのみ表示されます)
Databricks Databricks-Machine-Learning-Associate 認定試験の出題範囲:
| トピック | 出題範囲 |
|---|
| トピック 1 | - ML Workflows: The topic focuses on Exploratory Data Analysis, Feature Engineering, Training, Evaluation and Selection.
|
| トピック 2 | - Scaling ML Models: This topic covers Model Distribution and Ensembling Distribution.
|
| トピック 3 | - Spark ML: It discusses the concepts of Distributed ML. Moreover, this topic covers Spark ML Modeling APIs, Hyperopt, Pandas API, Pandas UDFs, and Function APIs.
|
| トピック 4 | - Databricks Machine Learning: It covers sub-topics of AutoML, Databricks Runtime, Feature Store, and MLflow.
|
参照:https://www.databricks.com/learn/certification/machine-learning-associate
弊社は無料Databricks Databricks-Machine-Learning-Associateサンプルを提供します
お客様は問題集を購入する時、問題集の質量を心配するかもしれませんが、我々はこのことを解決するために、お客様に無料Databricks-Machine-Learning-Associateサンプルを提供いたします。そうすると、お客様は購入する前にサンプルをダウンロードしてやってみることができます。君はこのDatabricks-Machine-Learning-Associate問題集は自分に適するかどうか判断して購入を決めることができます。
Databricks-Machine-Learning-Associate試験ツール:あなたの訓練に便利をもたらすために、あなたは自分のペースによって複数のパソコンで設置できます。
TopExamは君にDatabricks-Machine-Learning-Associateの問題集を提供して、あなたの試験への復習にヘルプを提供して、君に難しい専門知識を楽に勉強させます。TopExamは君の試験への合格を期待しています。
弊社は失敗したら全額で返金することを承諾します
我々は弊社のDatabricks-Machine-Learning-Associate問題集に自信を持っていますから、試験に失敗したら返金する承諾をします。我々のDatabricks Databricks-Machine-Learning-Associateを利用して君は試験に合格できると信じています。もし試験に失敗したら、我々は君の支払ったお金を君に全額で返して、君の試験の失敗する経済損失を減少します。
弊社のDatabricks Databricks-Machine-Learning-Associateを利用すれば試験に合格できます
弊社のDatabricks Databricks-Machine-Learning-Associateは専門家たちが長年の経験を通して最新のシラバスに従って研究し出した勉強資料です。弊社はDatabricks-Machine-Learning-Associate問題集の質問と答えが間違いないのを保証いたします。

この問題集は過去のデータから分析して作成されて、カバー率が高くて、受験者としてのあなたを助けて時間とお金を節約して試験に合格する通過率を高めます。我々の問題集は的中率が高くて、100%の合格率を保証します。我々の高質量のDatabricks Databricks-Machine-Learning-Associateを利用すれば、君は一回で試験に合格できます。
安全的な支払方式を利用しています
Credit Cardは今まで全世界の一番安全の支払方式です。少数の手続きの費用かかる必要がありますとはいえ、保障があります。お客様の利益を保障するために、弊社のDatabricks-Machine-Learning-Associate問題集は全部Credit Cardで支払われることができます。
領収書について:社名入りの領収書が必要な場合、メールで社名に記入していただき送信してください。弊社はPDF版の領収書を提供いたします。
一年間の無料更新サービスを提供します
君が弊社のDatabricks Databricks-Machine-Learning-Associateをご購入になってから、我々の承諾する一年間の更新サービスが無料で得られています。弊社の専門家たちは毎日更新状態を検査していますから、この一年間、更新されたら、弊社は更新されたDatabricks Databricks-Machine-Learning-Associateをお客様のメールアドレスにお送りいたします。だから、お客様はいつもタイムリーに更新の通知を受けることができます。我々は購入した一年間でお客様がずっと最新版のDatabricks Databricks-Machine-Learning-Associateを持っていることを保証します。