Kyle
10/15/2024, 5:06 AMDesmond Cheong
10/15/2024, 5:19 AMKyle
10/15/2024, 5:28 AMdef components(df: DataFrame) -> DataFrame:
b = df.select(daft.to_struct("u", "v").alias("e")).collect()
while True:
a = (b
.select(large_star_map(col("e.u"), col("e.v")).alias("e"))
.explode("e")
.select("e.*")
.groupby("u").agg_list("v")
.select(large_star_reduce(col("u"), col("v")).alias("e"))
.explode("e")
.where(~col("e").is_null())
.distinct()
.collect()
)
b = (a
.select(small_star_map(col("e.u"), col("e.v")).alias("e"))
.select("e.*")
.groupby("u").agg_list("v")
.select(small_star_reduce(col("u"), col("v")).alias("e"))
.explode("e")
.where(~col("e").is_null())
.distinct()
.collect()
)
# check convergence
a_hash = a.select(col("e").hash().alias("hash")).sum("hash").to_pydict()["hash"][0]
b_hash = b.select(col("e").hash().alias("hash")).sum("hash").to_pydict()["hash"][0]
if a_hash == b_hash:
return b.select("e.*")Kyle
10/15/2024, 5:29 AMjay
10/15/2024, 5:32 AMKyle
10/15/2024, 5:35 AMKyle
10/15/2024, 5:38 AMjay
10/15/2024, 5:39 AMjay
10/15/2024, 5:40 AMKyle
10/15/2024, 5:47 AMDesmond Cheong
10/15/2024, 5:52 AMKyle
10/15/2024, 5:56 AMdf = read_parquet()
df = ...
df.write_parquet(path)
df.read_parquet(path)
df2 = other_processes(df)
vs
df = read_parquet()
df = ...
df2 = other_processes(df)Kyle
10/15/2024, 6:02 AMdf = read_parquet()
df = ...
df.write_parquet(path)
df2 = other_processes(df)Desmond Cheong
10/15/2024, 6:04 AMdf = read_parquet()
df = ...
df.write_parquet(path)
df = read_parquet(path)
df2 = other_processes(df)
?Kyle
10/15/2024, 6:05 AMDesmond Cheong
10/15/2024, 6:06 AMKyle
10/15/2024, 6:06 AMKyle
10/15/2024, 6:06 AMDesmond Cheong
10/15/2024, 6:09 AMdef f():
df = read_parquet()
df = ...
df.write_parquet(path)
f()
df = read_parquet(path)
df2 = other_process(df)Desmond Cheong
10/15/2024, 6:13 AMA = object1
A = object2
The actual steps under the hood are:
1. Object1 will be created and assigned to dict['A']
2. Object2 is created
3. Dict['A'] is reassigned to point to object2
4. Object1 is dropped
So both objects (in our case both dataframes) will exist in memory during the reassignmentDesmond Cheong
10/15/2024, 6:14 AMKyle
10/15/2024, 6:16 AMDesmond Cheong
10/15/2024, 6:17 AMKyle
10/15/2024, 6:17 AMKyle
10/15/2024, 6:18 AMDesmond Cheong
10/15/2024, 6:19 AMDesmond Cheong
10/15/2024, 6:19 AMKyle
10/15/2024, 6:20 AMDesmond Cheong
10/15/2024, 6:21 AMDesmond Cheong
10/15/2024, 6:22 AMKyle
10/15/2024, 6:23 AMKyle
10/15/2024, 6:25 AMKyle
10/15/2024, 6:32 AM