Databricks Associate-Developer-Apache-Spark-3.5 試験概要:
| 認定ベンダー: | Databricks |
|---|---|
| 試験名: | Databricks Certified Associate Developer for Apache Spark 3.5 - Python |
| 試験番号: | Associate-Developer-Apache-Spark-3.5 |
| 試験時間: | 90 分 |
| 受験料: | $200 USD |
| 合格点: | 70% |
| 試験形式: | Multiple Choice |
| 認定の有効期間: | 2年間 |
| 対応言語: | English |
| 出題数: | 45 |
| 推奨トレーニング: | Databricks Academyによるトレーニング |
| 受験申し込み: | Databricks認定資格の登録 |
| サンプル問題: | Databricks Associate-Developer-Apache-Spark-3.5 サンプル問題 |
| 受験方法: | オンライン監督付きまたは会場での監督付き受験 |
| 前提条件: | 公式な受験資格は不要です。PySparkおよびDataFrame APIを使用した実務経験6か月以上が推奨されます |
| 公式シラバスのURL: | https://www.databricks.com/learn/certification/apache-spark-developer-associate |
Databricks Associate-Developer-Apache-Spark-3.5 試験シラバストピック:
| セクション | 比重 | 目標 |
|---|---|---|
| トピック 1: Structured Streaming | 10% | - ストリーミングクエリの定義 - 耐障害性と状態管理 - ストリーミングの概念とアーキテクチャ - 出力モードとトリガー |
| トピック 2: Apache Sparkのアーキテクチャとコンポーネント | 20% | - 実行モードとデプロイモード - Sparkアーキテクチャの概要 - シャッフル、アクション、ブロードキャスト - 耐障害性とガベージコレクション - 実行階層と遅延評価 |
| トピック 3: Apache Spark DataFrame APIアプリケーションのトラブルシューティングとチューニング | 10% | - 変換処理とアクションの最適化 - パフォーマンス上のボトルネックの特定 - メモリとリソース使用量の管理 - デバッグとログの取得 |
| トピック 4: Apache SparkにおけるPandas APIの利用 | 5% | - SparkにおけるPandas APIの概要 - 主な違いと制限事項 - Pandas構造とSpark構造の相互変換 |
| トピック 5: Spark Connectを使用したアプリケーションのデプロイ | 5% | - Spark Connect経由でのアプリケーション実行 - リモートSparkクラスターへの接続 - Spark Connectのアーキテクチャ |
| トピック 6: Apache Spark DataFrame APIを使用したアプリケーション開発 | 30% | - データのパーティショニングとバケッティング - 各種形式でのデータの読み書き - データセットの結合と統合 - データのフィルタリング、並べ替え、集計 - ユーザー定義関数(UDF) - 欠損値の処理とデータ品質の確保 - 列の選択、名前変更、変更 - DataFrameの作成とスキーマの定義 |
| トピック 7: Spark SQLの利用 | 20% | - 関数と式の操作 - DataFrameとSpark SQLの連携 - カタログおよびメタデータAPIの利用 - SQLクエリの実行 |
Databricks Certified Associate Developer for Apache Spark 3.5 - Python 認定 Associate-Developer-Apache-Spark-3.5 試験問題:
問題 #1
30 of 55.
A data engineer is working on a num_df DataFrame and has a Python UDF defined as:
def cube_func(val):
return val * val * val
Which code fragment registers and uses this UDF as a Spark SQL function to work with the DataFrame num_df?
A. spark.createDataFrame(cube_func("num")).show()
B. num_df.register("cube_func").select("num").show()
C. num_df.select(cube_func("num")).show()
D. spark.udf.register("cube_func", cube_func)
num_df.selectExpr("cube_func(num)").show()
問題 #2
44 of 55.
A data engineer is working on a real-time analytics pipeline using Spark Structured Streaming.
They want the system to process incoming data in micro-batches at a fixed interval of 5 seconds.
Which code snippet fulfills this requirement?
A. query = df.writeStream \
.outputMode("append") \
.trigger(processingTime="5 seconds") \
.start()
B. query = df.writeStream \
.outputMode("append") \
.trigger(continuous="5 seconds") \
.start()
C. query = df.writeStream \
.outputMode("append") \
.start()
D. query = df.writeStream \
.outputMode("append") \
.trigger(once=True) \
.start()
問題 #3
A data scientist wants each record in the DataFrame to contain:
The first attempt at the code does read the text files but each record contains a single line. This code is shown below:
The entire contents of a file
The full file path
The issue: reading line-by-line rather than full text per file.
Code:
corpus = spark.read.text("/datasets/raw_txt/*") \
.select('*', '_metadata.file_path')
Which change will ensure one record per file?
Options:
A. Add the option wholetext=False to the text() function
B. Add the option wholetext=True to the text() function
C. Add the option lineSep=", " to the text() function
D. Add the option lineSep='\n' to the text() function
問題 #4
A data engineer wants to create a Streaming DataFrame that reads from a Kafka topic called feed.
Which code fragment should be inserted in line 5 to meet the requirement?
Code context:
spark \
.readStream \
.format("kafka") \
.option("kafka.bootstrap.servers", "host1:port1,host2:port2") \
.[LINE 5] \
.load()
Options:
A. .option("kafka.topic", "feed")
B. .option("topic", "feed")
C. .option("subscribe", "feed")
D. .option("subscribe.topic", "feed")
問題 #5
48 of 55.
A data engineer needs to join multiple DataFrames and has written the following code:
from pyspark.sql.functions import broadcast
data1 = [(1, "A"), (2, "B")]
data2 = [(1, "X"), (2, "Y")]
data3 = [(1, "M"), (2, "N")]
df1 = spark.createDataFrame(data1, ["id", "val1"])
df2 = spark.createDataFrame(data2, ["id", "val2"])
df3 = spark.createDataFrame(data3, ["id", "val3"])
df_joined = df1.join(broadcast(df2), "id", "inner") \
.join(broadcast(df3), "id", "inner")
What will be the output of this code?
A. The code will result in an error because broadcast() must be called before the joins, not inline.
B. The code will work correctly and perform two broadcast joins simultaneously to join df1 with df2, and then the result with df3.
C. The code will fail because the second join condition (df2.id == df3.id) is incorrect.
D. The code will fail because only one broadcast join can be performed at a time.
解説:
| 問題 #1 正解: D | 問題 #2 正解: A | 問題 #3 正解: B | 問題 #4 正解: C | 問題 #5 正解: B |














1057 お客様のコメント
品質保証JPexamはIT認定試験のシラバスに従って、試験問題の範囲を正確に絞って、的中率が99%の最新問題集を捧げます。
1年間の無料更新サービスJPexamは1年以内に問題集の無料更新サービスを提供し、お客様がいつでも最新版の問題集を持つことを保証いたします。もし試験の内容が変更されたら、弊社は直ちにお客様にお知らせします。それに、弊社の問題集が更新されたら、早速メールで最新バージョンを送付いたします。
全額返金JPexamの問題集を利用すると、短時間で勉強しても試験に合格できるのを保証いたします。試験に不合格になってしまった場合、弊社は全額返金いたします。(
ご購入前のお試しJPexamは問題集のサンプルを無料で提供いたします。ご購入前にサンプルを試用して製品の品質を確認することができます。ご遠慮なく利用してください。
