dataset-viber (Dataset Viber)

davidberenstein1957

posted an update 3 months ago

Post

2657

🚨 Phare LLM benchmark V2: Reasoning models don't guarantee better security

Read the full blog here: https://huggingface.co/blog/davidberenstein1957/phare-llm-benchmark-v2

davidberenstein1957

posted an update 8 months ago

Post

621

Announcing RealPerformance, a dataset of functional issues of language models that mirrors failure patterns identified through rigorous testing in real LLM agents

https://huggingface.co/blog/davidberenstein1957/realperformance-llm-business-compliance

davidberenstein1957

posted an update 9 months ago

Post

412

🚨 LLMs recognise bias but also reproduce harmful stereotypes: an analysis of bias in leading LLMs

I've written a new entry in our series on the Giskard, BPIFrance and Google Deepmind Phare benchmark(phare.giskard.ai).

This time it covers bias: https://huggingface.co/blog/davidberenstein1957/llms-recognise-bias-but-also-produce-stereotypes

Previous entry on hallucinations: https://huggingface.co/blog/davidberenstein1957/phare-analysis-of-hallucination-in-leading-llms

1 reply

·

davidberenstein1957

posted an update 10 months ago

Post

1771

I created a collection of FLUX.1 models but 4x faster PrunaAI/flux1-but-4x-faster-66c0b7340836dd7a55e9c0ea

davidberenstein1957

posted an update 11 months ago

Post

1100

Good answers are not necessarily factual answers: an analysis of hallucination in leading LLMs

https://huggingface.co/blog/davidberenstein1957/phare-analysis-of-hallucination-in-leading-llms

davidberenstein1957

updated a model 11 months ago

dataset-viber/stable-diffusion-v1-4-smashed

Text-to-Image • Updated May 6, 2025 • 1

davidberenstein1957

published a model 11 months ago

dataset-viber/stable-diffusion-v1-4-smashed

Text-to-Image • Updated May 6, 2025 • 1

davidberenstein1957

posted an update 11 months ago

Post

2280

🔥 Announcing FLUX-Juiced: The Fastest Image Generation Endpoint (2.6x faster)!

Optimisations are widely applied and can reduce inference time, but their impact on quality often remains unclear, so we decided to challenge the status quo and create our own optimised version of FLUX.1[dev] called FLUX-juiced.

Blog: https://huggingface.co/blog/PrunaAI/flux-fastest-image-generation-endpoint

davidberenstein1957

posted an update 11 months ago

Post

1745

🧑‍🏫 I wrote a brief blogpost to give An Introduction to AI Model Optimization Techniques!

URL: https://huggingface.co/blog/PrunaAI/introduction-to-ai-model-optimization-techniques

davidberenstein1957

posted an update 12 months ago

Post

1411

RealHarm: A Collection of Real-World Language Model Application Failure

I'm David from Giskard, and we work on securing your Agents.
Today, we are launching RealHarm: a dataset of real-world problematic interactions with AI agents, drawn from publicly reported incidents.

Check out the dataset and paper: https://realharm.giskard.ai/

davidberenstein1957

posted an update about 1 year ago

Post

2114

🚨 New Bonus Unit: Tracing & Evaluating Your Agent! 🚨

Learn how to transform your agent from a simple demo into a robust, reliable product ready for real users.

UNIT: https://huggingface.co/learn/agents-course/bonus-unit2/introduction

In this unit, you'll learn:
- Offline Evaluation – Benchmark and iterate your agent using datasets.
- Online Evaluation – Continuously track key metrics such as latency, costs, and user feedback.

Happy testing and improving!

Thanks Langfuse team!

davidberenstein1957

posted an update about 1 year ago

Post

2439

🔥 Text2SQL, explore and share any data analysis!

🤗 Hugging Face - Dataset Studio is an amazing new feature.

🚀 Start yourself: https://huggingface.co/datasets/fka/awesome-chatgpt-prompts/viewer/default/train

📺 YouTube: https://youtu.be/5LUZq7MHolA?feature=shared

davidberenstein1957

posted an update about 1 year ago

Post

4269

🥊 Epic Agent Framework Showdown! Available today!

🔵 In the blue corner, the versatile challenger with a proven track record of knowledge retrieval: LlamaIndex!

🛑 In the red corner, the defender, weighing in with lightweight efficiency: Hugging Face smolagents!

🔗 URL:

agents-course

We just published the LlamaIndex unit for the agents course, and it is set to offer a great contrast between the smolagents unit by looking at

- What makes llama-index stand-out
- How the LlamaHub is used for integrations
- Creating QueryEngine components
- Using agents and tools
- Agentic and multi-agent workflows

The team has been working flat-out on this for a few weeks. Supported by Logan Markewich and Laurie Voss over at LlamaIndex.

Who won? You decide!

davidberenstein1957

posted an update about 1 year ago

Post

3059

🫸 New release to push vector search to the Hub with vicinity and work with any serialisable objects.

🧑‍🏫 KNN, HNSW, USEARCH, ANNOY, PYNNDESCENT, FAISS, and VOYAGER.

🔗 Example Repo: minishlab/my-vicinity-repo

davidberenstein1957

posted an update about 1 year ago

Post

3329

🚀 Find banger tools for your smolagents!

I created the Tools gallery, which makes tools specifically developed by/for smolagents searchable and visible. This will help with:
- inspiration
- best practices
- finding cool tools

Space: davidberenstein1957/smolagents-and-tools

1 reply

·

davidberenstein1957

posted an update about 1 year ago

Post

2494

Fine-tune Deepseek-R1 with a Synthetic Reasoning Dataset

Blog: https://huggingface.co/blog/sdiazlor/fine-tune-deepseek-with-a-synthetic-reasoning-data

davidberenstein1957

posted an update about 1 year ago

Post

2089

Agentic RAG: Applied, visual, and step-by-step! 🐾

Get familiar with the Agents and tools, not the bells and whistles!

Retrieve - Augment and now GENERATE.

part 3: https://huggingface.co/blog/davidberenstein1957/ai-blueprint-agentic-rag-part-3-generate

davidberenstein1957

posted an update about 1 year ago

Post

2895

Anyone can create free hosted tools for their AI agents! 🔥

Agentic RAG stack part 2 - augment
Augment retrieval results by reranking optimises content without increasing time too much

part2: https://huggingface.co/blog/davidberenstein1957/ai-blueprint-agentic-rag-part-2-augment
code: https://github.com/huggingface/ai-blueprint

davidberenstein1957

posted an update about 1 year ago

Post

2026

Creating an agentic RAG stack on the Hugging Face Hub - part 1 - retrieval (1/5).

🚀 Web apps and microservices included!

Chunk, embed and index documents at a huge scale without overhead.

Blog: https://huggingface.co/blog/davidberenstein1957/ai-blueprint-agentic-rag-part-1-retrieve

davidberenstein1957

posted an update about 1 year ago

Post

1645

tldr; Parquet is awesome, DuckDB too!

Datasets on the Hugging Face Hub rely on parquet files. We can interact with these files using DuckDB as a fast in-memory database system. One of DuckDB’s features is vector similarity search which can be used with or without an index.

blog:
https://huggingface.co/learn/cookbook/vector_search_with_hub_as_backend

AI & ML interests

Team members 1

dataset-viber's activity