Kyle
10/11/2024, 5:56 AMDesmond Cheong
10/11/2024, 6:00 AMdocs/source/api_docs/expressions.rst?Kyle
10/11/2024, 6:06 AMKyle
10/11/2024, 6:31 AM@daft.udf(return_dtype=daft.DataType.struct({"top_key":daft.DataType.string(), "max_value": daft.DataType.uint64()}))
def get_most_frequent_key(count_series):
counts_list = count_series.to_pylist()
results = []
for counts in counts_list:
top_key = None
max_value = 0
for count in counts:
key, value = count
if value > max_value:
top_key = key
max_value = value
results.append({"top_key": top_key, "max_value": max_value})
return resultsDesmond Cheong
10/11/2024, 6:43 AMlist.sort , but it only takes primitive types and does not support custom predicates)
with the two pieces above, we could naturally sort the value counts in descending order of counts, then just call list.get on the first element to achieve what you wantKyle
10/11/2024, 6:47 AMDesmond Cheong
10/11/2024, 7:00 AMmap_entries(map) that gives a list of (key, values), but then you can't really sort this after (that I'm aware of). Spark has map_entries and array_sort , but for value_counts you'd have to 1. get distinct elements for each list, 2. roll your own lambda function in a transform to aggregate the counts for each distinct value.Andrew Gazelka
10/11/2024, 9:22 AM