Hey guys, One thing I've been curious about and me...
# general
c
Hey guys, One thing I've been curious about and meaning to ask is what the expected workflow for generating simulated data for the sake of validating pipelines. If I want to create a Daft series with a fixed length, suppose with random int values as a placeholder, I have been using a UDF that receives a series as input, and generates the random values using the numpy and outputing an entire series, like this:
Copy code
@daft.udf(return_dtype=daft.DataType.python())
def make_num_points(x: daft.Series):
    """Simulate column type and structure."""
    rng = np.random.default_rng()
    return [rng.integers(low=0, high=10, size=len(n)) for n in x.to_pylist()]
Is there a recommended way to take this concept in Daft to a 1-2-liner that can be quickly used to prototype pipelines and generate the data JIT-style? There are several other use cases like subtle differences in Optional vs None column inputs, and how we can parameterize existing input spaces most efficiently without having to use a large column that may not be needed. Has anyone else encountered these use cases?
r
Hey Cora, if I understand correctly, your current implementation (using the UDF with numpy.random inside) works, but you are wondering if there's a more standardized way of doing it?
c
Yes!
r
Could you highlight what your use-case is? Chatted with @Kevin Wang about this and he mentioned that you could try generating a list of random values and then calling
daft.from_pydict
on it. I guess the more "suggested" workflow here would be more dependent on what your needs are out of it.
c
https://github.com/Eventual-Inc/Daft/blob/b87e0a38d0e324351446ef136703bc4e03b702b2/daft/io/_generator.py#L20 We do have a non-public API that creates dataframes directly from python functions, useful if each function is parametrized differently to return different variations in data, or if you want to test with multiple partitions