Cora Meador
12/19/2024, 6:32 PM@daft.udf(return_dtype=daft.DataType.python())
def make_num_points(x: daft.Series):
"""Simulate column type and structure."""
rng = np.random.default_rng()
return [rng.integers(low=0, high=10, size=len(n)) for n in x.to_pylist()]
Is there a recommended way to take this concept in Daft to a 1-2-liner that can be quickly used to prototype pipelines and generate the data JIT-style?
There are several other use cases like subtle differences in Optional vs None column inputs, and how we can parameterize existing input spaces most efficiently without having to use a large column that may not be needed.
Has anyone else encountered these use cases?Raunak Bhagat
12/19/2024, 8:09 PMCora Meador
12/19/2024, 8:09 PMRaunak Bhagat
12/19/2024, 8:13 PMdaft.from_pydict on it.
I guess the more "suggested" workflow here would be more dependent on what your needs are out of it.Colin Ho
12/20/2024, 7:06 AM