Bloom Filters Explained: The Probabilistic Data Structure Powering Instagram, Google, and High-Scale Systems

Iniciado por joomlamz, Hoje at 14:15

Respostas: 1   |   Visualizações: 8

Tópico anterior - Tópico seguinte

0 Membros e 1 Visitante estão a ver este tópico.

Saudações, comunidade do **webmastersmz.com**!

Como especialista em tecnologia, analisei recentemente o fascinante tópico **"Playwright Agents: The Architecture of Self-Healing"** (Agentes Playwright: A Arquitetura de Autocura). Este tema está a dar muito que falar no ecossistema de desenvolvimento e testes de software, trazendo inovações profundas para a área de automação web.

De seguida, destaco os pontos técnicos principais discutidos no artigo original:

1. **O Problema dos Testes Frágeis (Flaky Tests):** Tradicionalmente, os testes automatizados com ferramentas como o Playwright dependem de seletores CSS ou XPath rígidos. Qualquer pequena alteração na interface do utilizador (UI) — como mudar o ID de um botão ou a estrutura de uma *div* — quebra o teste, exigindo manutenção manual constante por parte dos programadores.
2. **Introdução dos Agentes de Autocura (Self-Healing):** A arquitetura apresentada combina o poder do Playwright com modelos de Inteligência Artificial (LLMs) e agentes autónomos. Quando um seletor falha durante a execução do teste, o agente não desiste; em vez disso, ele analisa o DOM atual, compreende a intenção original do teste e recalcula dinamicamente um novo seletor válido em tempo de execução.
3. **Redução da Dívida Técnica:** Com esta arquitetura, a manutenção de testes deixa de ser uma tarefa reativa e exaustiva. Os agentes aprendem com as mudanças na interface, permitindo que as equipas de desenvolvimento foquem a sua energia na criação de novas funcionalidades em vez de corrigirem testes quebrados.
4. **Resiliência e Escalabilidade:** A integração de IA diretamente no ciclo de testes abre portas para uma automação muito mais robusta, capaz de antecipar comportamentos e garantir uma entrega contínua (CI/CD) muito mais estável.

**O que acham desta abordagem, caros colegas do webmastersmz.com?**
Já tiveram a oportunidade de implementar agentes baseados em IA nos vossos fluxos de trabalho ou ainda preferem os métodos tradicionais de seletores estáticos? Como vêem o impacto disto na produtividade das nossas equipas de desenvolvimento em Moçambique e no mundo? **Deixem as vossas opiniões e experiências nos comentários abaixo e vamos enriquecer este debate!**

---

Para garantir que os vossos projetos e fóruns rodam sem falhas, convido-vos a conhecer as soluções de alojamento de alta performance da AplicHost em [https://aplichost.com](https://aplichost.com).


                     Bloom Filters Explained: The Probabilistic Data Structure Powering Instagram, Google, and High-Scale Systems
               




Tópico:
                     Bloom Filters Explained: The Probabilistic Data Structure Powering Instagram, Google, and High-Scale Systems
               
Categoria: Tutoriais | FreeCodeCamp Premium
Idioma Principal: Português (Conteúdo de Tecnologia)

Conteúdo do Tutorial / Guia Passo a Passo:
-------------------------------------------------------------------------
Instagram has over 500 million registered usernames. When a new user tries to register, the platform needs to answer one question almost instantly: has this username already been taken?

The naïve answer is a database query. Pull the users table, search for the username, return whether it exists. This works fine at small scale. But at 500 million records, queried millions of times every day, it becomes a serious bottleneck. Even with indexing, the database is doing expensive work for every single registration attempt.

Luckily, there's a smarter approach. Before the database is ever touched, you ask a different system a much faster question. That system gives you one of two answers:

Definitely not here: This is guaranteed. Zero exceptions. The username is available and you can skip the database entirely.

Probably here: This is not guaranteed. The username might be taken, or this might be a false alarm. You need to confirm with the database.

That system is a Bloom Filter. It can't tell you with certainty that something exists. But it can tell you with absolute certainty that something does not exist. And in systems at scale, that one-sided guarantee eliminates the vast majority of expensive database queries.

Table of Contents

• Prerequisites

• What is a Bloom Filter?

• The Two Components of a Bloom Filter

• How Adding an Item Works

• How Checking an Item Works

• False Positives Explained

• Why You Can't Delete From a Bloom Filter

• The False Positive Rate

• Implementation in Dart

• Implementation in C#

• Where Bloom Filters Are Used in Real Systems

• When to Use a Bloom Filter

• When Not to Use It

• Conclusion

Prerequisites

Before reading this article, you should be comfortable with:

• Basic data structures: arrays and how indexing works

• What a hash function is at a conceptual level: a function that takes an input and produces a fixed-size output

• Basic programming concepts in either Dart or C#

You don't need to understand database internals or distributed systems deeply. Where these appear in this article, they serve only to illustrate why Bloom Filters matter in real engineering.

What is a Bloom Filter?

A Bloom Filter is a probabilistic data structure that represents a set without storing the actual values in that set.

The word probabilistic is the important one. Unlike a regular set or a database, a Bloom Filter doesn't give you definitive answers about membership. It gives you probabilistic answers. And the probability is deliberately asymmetric:

It will never tell you something is absent when it's actually present. This is called having no false negatives.

It might occasionally tell you something is present when it is actually absent. This is called a false positive. The rate at which this happens is small, controlled, and mathematically predictable.

This asymmetry is what makes Bloom Filters useful. The "definitely not here" answer is trustworthy. The "probably here" answer is a signal to go confirm elsewhere.

The Two Components of a Bloom Filter

A Bloom Filter is built from two things.

1. A bit array
A fixed-size array of bits, all initialized to zero. This is the entire storage of the filter. Not strings or objects, just bits. Zeros and ones.

Position: 0  1  2  3  4  5  6  7  8  9  10 11 12 13 14 15
Value:     0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0

The size of this array i

... [O tutorial continua no link abaixo] ...


Joomlamz
Consultoria em Informática
-------------------------------------------------------
Especialista em Sistemas Web & Manutenção de Servidores.
A desenvolver o novo AplPortal com suporte a PHP 8.
Precisa de ajuda profissional? Contacte-me.

Tags: