Rendered at 23:49:10 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
muti 21 hours ago [-]
Weird how the readme talks about hash collisions, but not in the way I would expect. Two blocked domains with the same hash isn't a problem, they both need to be blocked.
Where hash collisions matter is false positives, e.g. if hash(google.com) = hash(adserver.com). There does appear to be a web dashboard and /unblock api so should be straightforward to resolve.
apefulsin 12 hours ago [-]
AI wrote the readme so it's unsurprising
rdsubhas 7 hours ago [-]
For all we know an AI bot wrote the above comment as well.
This is a worrying trend on HN. Nearly for all submissions, there is one unquantified, unproven comment at the top saying it's AI, with just emotion and no proof. This is the new witch hunt, or karma farming without having to contribute anything positive to the discussion.
ac29 6 hours ago [-]
> Nearly for all submissions, there is one unquantified, unproven comment at the top saying it's AI, with just emotion and no proof
nearly all of the commits in the repo are co-authored by claude
Muhammad523 3 hours ago [-]
It is proven because Claude is listed as a contributor. Its not proof-less and completely emotion based
apefulsin 6 hours ago [-]
you could just go straight to the source instead of speculating. read the readme. it's obviously AI.
zamadatix 20 hours ago [-]
Yeah, I think they have it backwards. As you say, regardless of how many entries you have locally, the collision risk comes from false positive hash matches not from worrying if the positive hashes collide.
4gotunameagain 13 hours ago [-]
I think they have it right.
If you have N entries where N >> 1, the probability of an arbitrary value colliding with the existing ones is Pcol(N+1) which is approx Pcol(N).
zamadatix 11 hours ago [-]
Edit: Skip to 4gotunameagain's reply, this probably doesn't contain anything helpful to read.
The math is correct, but it's the math for the question:
While the thing they probably want to check the collision math for is:
blockHashes.has(hash(domainNotInList))
As that's the risk you do the wrong action due to a hash collision. Importantly, the 2nd calculation is not going to be bounded by the size of your block list, it's going to be bounded by the number of domains overall.
The former question isn't really useful - you can have a 100 entry blocklist with 100 collisions and it still works the same as if it had 0 collision since the action on all 100 is the same. That's more a problem for generating unique IDs per hash whereas this is flipped around because it's about classification.
4gotunameagain 11 hours ago [-]
No, it is for the right (second) question.
If we assume uniform sampling for both the blocklist (size N) and the domain to be visited (not in the block list), the collision probability we care about is Pcol(N+1).
The number of existing domains does not matter, only the size of the block list.
Of course when we consider all 401 million domains there will be many more collisions, but in each case we only care about a collision in N+1
zamadatix 11 hours ago [-]
Ah, I get what you're saying to do with the collision probability and that should definitely make sense but whatever they actually did in the readme ended up a factor of ~4 off the result that approach should give. E.g. the much simpler direct path of 2^40/537000 for that "next domain" question is giving me 1 in ~2.05 million rather than 1 in ~537k.
Edit: Or maybe they formulated that path to derivation and how to map it back but couldn't quickly find a tool able to approximate the birthday problem to that scale so they tried direct testing instead? If, e.g., they were slowly upping the values to see when a collision occurred it would make sense they ran into a collision earlier than would be expected as they (effectively) gave more than 1 trial. Or just random chance too if that's really the path they took I suppose :).
Edit2: After looking at the readme for far too long, I saw the original readme was in Japanese. Running that through a good translator actually clarifies or corrects a lot of the wording like saying "In the latter case, one other unlucky domain also gets blocked" rather than focusing on the number of domains over-blocked out of the 537000 only. With this version of the text I think you're definitely right, the math approach used should have given them the right number as the original text is already flipping things back to what happens with the one new query but they (apparently) just didn't do the actual calculation with it.
Sharp catch :).
4gotunameagain 10 hours ago [-]
You don't need to simulate, it is a well studied problem [1]
It seems that the average number of expected collisions for 537k would be 0.13 and not 1:
the average value of pairs of individuals with the same birthday (or hash collisions) is E[X] = k*(k-1)/(2*n). Plugging in k=2^40 for 40 bits and n=537e3 I get their result of 0.1311
Thanks for the nerdsnipe, I ran out of my daily LLM mistake fixing quota !
I don't think this framing in this thread lines up with what the readme says. Also note that the python script to build the blocklist [1] reports the number of collisions found when generating the hashes.
I read their statement as "there were 0 collisions when building a blocklist for 141k domains, and 1 collision for a blocklist of 537k domains", and the script points at that number coming directly from collisions within the blocklist.
> Two blocked domains with the same hash isn't a problem, they both need to be blocked.
This is a problem if the second one is actually a major useful domain. If HN and adserver have the same hash, you have a problem. This could be described as a false positive for HN.
> Where hash collisions matter is false positives, e.g. if hash(google.com) = hash(adserver.com)
This sounds like the same issue I describe above where you need to use Google but block adserver. Otherwise it's only a problem if you implemented allow lists, right? This is when you explicitly allow Google and implicitly allow adserver along with it because of the hash collision.
muti 12 minutes ago [-]
Sorry was a bit unclear, by two blocked domains I meant two domains in the original blocklist having the same hash, i.e. hn and adserver are both in our desired blocklist in your example
yoavm 16 hours ago [-]
If you're ever thinking about getting an ESP32-C3, do yourself a favor and get the variant that you can connect an external antenna to. The normal C3 has a built-in antenna that is extremely weak, making it useless for most things I was planning using it for.
z2 11 hours ago [-]
And make sure to read the reviews as there are apparently fraudulent versions out there that don't include the 4MB flash chip. That said I took a gamble on a bunch of $2 C3 super minis, and they've had surprisingly adequate signal strength for indoor applications, even in the basement. (-60's dBm)
n8henrie 14 hours ago [-]
You can find instructions for making and soldering on a small antenna to the common "mini" dev board that reportedly helps quite a bit
utsavmishra25 12 hours ago [-]
yes and its such a important detail to have in later ESP projects. Having the ability to add antennas onto ESP-32 is a blessing-in-disguise that helps with more complex projects at a range later on
RicoElectrico 13 hours ago [-]
You mean chip antenna (one that looks like a big SMD resistor)? Yes, they're garbage. But this is orthogonal to the chip itself.
Also, PCB antennas are reasonably good for most indoor situations.
BLKNSLVR 18 hours ago [-]
I'm a bit of a paranoid freak that likes lists, so I've got a PiHole that has a total list of 14-16M blocked domains.
Great idea, and good for casual blocking, but I'm almost moving to an "allow list" mindset. This solution would probably work better for that, I wonder if the good parts of the internet would fit into an 140k list.
reader9274 17 hours ago [-]
Do you host your lists somewhere you can share?
BLKNSLVR 14 hours ago [-]
I'm currently (very slowly, in the small gaps between 'life') setting up a site the host the lists and also explains them and their usage.
Since it's not ready, it'll be best if I point you to the sources I've aggregated from:
I've aggregated the lists into four tiers: Base, Recommended, Aggressive, and Paranoid. And I've got two sets of these, one that includes Newly Registered Domains (Full directory), and one that doesn't include NRDs (NoNRD directory, because NRDs are ... heavy: Paranoid list with NRDs: 9.1 million records, without NRDs: 3.7 million records).
Aggregation of lists under the topic that matches the file name. Fake News list needs to be updated as it contains a _way_ overbroad list that I need to remove from the aggregation.
genxy 8 hours ago [-]
This information would make a great neocities page.
1vuio0pswjnm7 5 hours ago [-]
"This solution would probably work better for that, I wonder if the good parts of the internet [read: www] would fit into a 140k list."
For almost two decades I have been using an "allowlist", in the beginning it was via authoritative DNS. When proxy software began supporting mapping the DNS data into memory, I eliminates the need to use DNS
I have enough historical DNS data to assess own needs and it's much smaller than 140k
The popular blocklists can be quite strange if one takes the time to peruse them
They include what appear to be some very "bad parts of internet (www)" that I doubt most users would ever visit. There is more in these files than only common ad and tracking domains. As such one could call them "bloated"
It begs the question of how these blocklist maintainers even know about these bad parts
IMHO, the greatest exposure to ubiquitous ads, tracking and telemetry is through use of third party resolvers (in real-time, immediately preceding associated HTTP requests), where some process on the user's computer can potentially look up _any_ domain
Pretending that one can anticipate every possible undesired lookup, e.g., via a "blocklist", is IMHO a high risk approach
But it's popular. To each his own
Every user is different. Relying on a "blocklist" that attempts to anticipate every possible undesired domain is not for me. By contrast, the "allowlist" approach has served me reliably for many years. It sometimes suprises me how well it continues to work year after year
Philosophically, IMO, the contrast between "allowlist versus blocklist" is the difference between (a) "going after what one wants" from the www versus (b) "avoiding what one doesn't want"
The quantity of (a) can be relatively small and manageable. But (b) only grows larger over time, a tsunami full of garbage
I've always thought "ad blocking" was a misnomer. Instead of "blocking" what we as users want to do is "not send requests" to ad/tracking/telemetry servers
The www (HTTP) works via requests and responses. Generally, users make requests and publishers (including advertisers) receive the requests and (usually) send responses. Users are generally not receiving requests and few users would consciously send requests (including POSTing their private data) to ad, tracking or telemetry servers. It is "developers" who are making these requests for ads/tracking/telemetry, not users, by automatically generating them with the software they distribute to users (including Javascripts)
IMO, we don't "block" ads because we never requested them to begin with. What we do is ensure that the requests _made by developers_ do not succeed, we "block" requests that we never made
Segue into debate about user "agency"...
oso2k 16 hours ago [-]
Tbh, for the ESP32-C3’s perf, an allow list would likely save RAM and CPU time.
IgnaciusMonk 9 hours ago [-]
use ublock regardless, lists wont remove everything. you can remove piholelist items from ublock (firefox one, not chrome one).
and every linux user should turn firewall on. and better download CISA hardening guides /best practices . or similar agency guides in countries other than US. they ARE free and mostly unversal, even if guide deals with other distro same principles apply to any system, just use AI to find right commands on your system.
99.999% of linux users who "converted to light of linux" based on nonsensical persuasion from perverts from LTT and other similar perverts, do not follow basic security practices. And these LTT perverts know this and they are laughing at your faces.
Windows in default configuration
(as is installed without users turning off updates and other stuff off )
is 1000 milion times safer than linux for DESKTOP in default installed state. Which is state in which 99.9999% of loooosers sorry i meant to say users are using it in.. And im not even talking about totally irrational things to install like Crossover, Lutris and other nonsense based on WINE. I have 80 calls per month about infostealers, spyware in general which stem from using these nonsensical solutions full of old frameworsk (VB, .net etc ) and no security at all. Even official games, on EVERY platform, are scanning whole networks, combing thru your stuff (documents, pictures, webbrowsing activity... ) and noone cares, this is madness.
Distro maintainers should pull their heads from their asses too, how it is possibel to have firewall not enabled in OS in 2026?? HOOOOW. Not even talking about other stuff than like K12 stuff like firewall.
waysa 16 hours ago [-]
I think it could fit even more domains using a Bloom Filter or similar probabilistic data structure. With a chance of false-positives of course. But that's a trade-off the project already makes.
1vuio0pswjnm7 20 hours ago [-]
Whitelist/allowlist is easier, e.g., it's smaller
Depends on the user but not everyone is visiting new websites everyday
Even for those that are, the number of domain-IP mappings needed will be relatively small
Definitely under 140,000
Most DNS data I use is "static", it rarely changes. As such most times I don't have to make DNS queries. I store the domain-IP mappings in proxy memory; this is faster than DNS
No "blocklist" needed
sheept 19 hours ago [-]
I would think that a regular user of Hacker News would be visiting new websites every day (though it'd definitely still be below 140k)
Etheryte 17 hours ago [-]
This is all a guesstimate, but my gut feel is that a considerable part of the HN population doesn't even read the linked content, only the comment section here.
jasonjmcghee 17 hours ago [-]
I sure hope that's not true. Maybe hit the comments section first?
dwedge 16 hours ago [-]
With github being the exception, I use comments to see if the article is AI. If it is, I prefer the condensed opinions in the comments. Not to say it can't be interesting I just don't want to waste time reading overly verbose generated text.
lhoff 16 hours ago [-]
Depends on the content, for all of these model release marketing sites, I for example only read the comments.
If i would estimate it, I only take a look at 1/4 of the links where i read the comments.
19 hours ago [-]
calgoo 13 hours ago [-]
Just make a captive portal that shows up on your screen for new domains where you can just click "add to allow" or "allow this once" or "add to deny". That way you build the dataset, like most of these tools its annoying in the beginning while you generate the dataset, but after a while its only a few pages a week.
timvdalen 17 hours ago [-]
> The trick everyone misses:
Please just write the first sentence of your README yourself
WithinReason 13 hours ago [-]
The trick everyone misses: binary search of hashes can be done in log(log(n))
dwedge 16 hours ago [-]
Apparently solving blocking extra domains due to hash collisions (reducing from 1 to 0) would be extra space "to solve a problem I don't have".
I dislike when LLMs talk to me like that. I hate it when humans do it, confidently spewing their overconfident llm assumptions to others
lifeisloving 16 hours ago [-]
I actually thought this was one of the more communicative parts of the readme and quite liked their explanation.
13 hours ago [-]
anilakar 17 hours ago [-]
There's no point in using PlatformIO for ESP32 MCUs. The native ESP-IDF extension works much better.
I would only recommend using it instead of the standard Arduino Processing IDE.
ricardobeat 12 hours ago [-]
I use arduino-cli exclusively, lighter than ESP-IDF and works completely standalone from the IDE.
Muhammad523 3 hours ago [-]
I was exited to read about this until I saw "Claude" listed as a contributor.
bilekas 13 hours ago [-]
Nice project and poc maybe, but I'm not seeing the practical application of this in a real world env, surely the latency is a deal breaker ?
fwip 21 hours ago [-]
Cool idea, latency might be too high, wish the docs weren't all AI vomit.
snailmailman 19 hours ago [-]
The biggest benefit of my local dns server is latency.
On wired internet, my dns is <1ms from my PC.
Upstream dns for me is pretty quick. Google and cloudflare dns are ~5ms from me.
But WiFi latency alone is ~8ms most of the time in my experience. On my fiber internet, pinging a dns server in some random upstream server miles away is lower latency than WiFi 10 feet away. But the real issue on WiFi is any packet loss at all adding 50-100ms to that at random depending on interference.
With DNS you are paying this latency cost all the time on nearly every request.
ktm5j 10 hours ago [-]
But this is an ESP.. it's a 160 MHz microcontroller that does not have wired ethernet, only wifi. Also it's serving data from a USB flash drive instead of from RAM. Latency is going to be through the roof.
fwip 6 hours ago [-]
I haven't looked at the math - I wonder if a probabilistic data structure for this dataset (e.g: bloom filter) could be small enough to fit in RAM and still be useful. On my home network, approximately 10-20% of queries hit our blocklist - getting a quick exit on even half of the ~80% of good queries might be noticeable in aggregate.
QuantumNomad_ 15 hours ago [-]
> With DNS you are paying this latency cost all the time on nearly every request.
Surely not? macOS, Windows, and I think most of the big Linux distros, all cache DNS responses. Probably the web browser itself does too.
fwip 10 hours ago [-]
Android and iOS do as well.
genxy 8 hours ago [-]
real/paying/cost/nearly
The real cost of using too much AI is nearly writing like them all the time.
snailmailman 7 hours ago [-]
I seldomly use AI. I, a human, wrote the comment.
thenthenthen 20 hours ago [-]
This will be slow for one user, let alone more than one
danw1979 17 hours ago [-]
“The trick everyone misses”
rubyfan 12 hours ago [-]
I really like this conceptually. However, multiple elements feel like slop and that’s throwing warning signs to me. The Claude readme file and the tiktok-ish fast cut youtube video felt like Vince the Sham-wow guy put them together. Ironically it feels like the product of the problem we’re trying to avoid.
yeah but it doesn’t replace the remote resolver, in most setups PiHole is another layer on top of the remote resolver
almog 12 hours ago [-]
But a remote resolver, such as the ones on the list, is not used to determine whether an item is contained in a set (0.0.0.0) but rather to retrieve the current IP of that domain.
The way Pi-Hole is used is to first determine whether a DNS should be blocked and only if it's not in that set, forward it to a remote resolver.
I'd think the local resolver should be at least an order of magnitude less than the remote to justify it.
jwillmer 15 hours ago [-]
pinhole is great but if the device fails your network is down. I swapped to nextdns because of that. maybe support of multiple devices would fix it - now that is this cheap it would not matter
.
import 12 hours ago [-]
You can have a 2 adguard home + adguard-sync. Works smoothly.
b3lvedere 12 hours ago [-]
But what if both fail? :)
close04 15 hours ago [-]
The simple answer is have 2. If the cost per unit is low enough than deploying 2 PiHoles is trivial.
The developers could make users' lives easier by implementing a clustering option that syncs the 2 out of the box. For static deployments, where you don't change the config too often, it's still decent as it is.
krtkush 14 hours ago [-]
I did the same (until my RPi zero stopped working). Both the devices share same vertical IP address and if one ever went down, the other would take over without any delay.
fwip 10 hours ago [-]
I have my router set up to send DNS requests to my adguard home server, and if there's no response, it falls back to one of Cloudflare's. So it fails-open, letting the ads in along with normal requests.
19 hours ago [-]
swiftcoder 8 hours ago [-]
Slop aside, I think this angle of hyper-optimising Adblock is interesting. I recently migrated my pihole from an intel minipc down to an OrangePI Zero 3. Saved ~5 watts of power, which lets me extend the battery backup duration of my network closet during a power outage.
aftbit 7 hours ago [-]
I'm weird but I need to use a pretty fast PC for a router anyway to get 2.5Gb with QoS and Wireguard for when I'm a road warrior, so I end up just running everything on there rather than having a separate machine for things like Pihole.
I really do need to clean up my DNS setup a bit though. It's pretty wacky right now.
peter_d_sherman 12 hours ago [-]
>"The trick everyone misses: you don't need to keep the blocklist in RAM.
Store the domains as sorted 40-bit hashes in flash and binary-search them.
140,000+ domains fit in ~0.7 MB of flash..."
Brilliant! (Although, what about potential hash collisions with domain names that shouldn't be blocked? Those are probably few and far between, so the utility of the block overall far outweighs any few false positives that might arise -- so I reiterate my claim of "brilliant"!)
Also (if you're a compression nerd like I am), it may be possible to analyze this table of 40-bit hashes further to find the longest binary substring that repeats the most across all of them (or binary substrings that just repeat a whole lot), and design a Huffman-like (or other tree or tree-like) data structure around those, and possibly compress the data further...
Of course, then the simplicity and elegance of the Binary Search might be lost, but it might be an interesting optimization exercise to know just how far that data could be compressed... "Hey Claude, try to optimize the compression of this data using alternative compression data structures, and report back to me!" (Or something like that! You know, get the AI coding agents to try different things -- or attempt to code it on your own (always better for understanding!), etc., etc.!)
Anyway, brilliant!
nicman23 18 hours ago [-]
i dont get pihole. just use a proper dns? if you want local, just use unbound?
Tajnymag 18 hours ago [-]
Pihole and unbound can work together. Pihole isn't a good DNS server, it's a good adblocking DNS server.
dannyw 17 hours ago [-]
people like ease of use, UI, docs, and community.
after all, why use opnSense; just use suricata and $PROPER_X! why use TrueNAS; just use $DISTRO and ZFS!
nicman23 15 hours ago [-]
it is easier to add adguard dns if you want ease of use
rekoil 17 hours ago [-]
"They were too preoccupied with whether they could, they never stopped to think whether they should"
Jokes aside, very impressive that it works!
manlymuppet 18 hours ago [-]
Wow, this is an atrocious name for a project haha. Cool though.
shevy-java 13 hours ago [-]
I have a weird opinion: I believe adblocking must become a basic human right. Conversely Google trying to deny ublock origin and other extensions, should be fined many billions of euros as punishment. Then be disbanded if unable to pay. That money also should directly go back to The People, since they suffered from this evil malpractice of Google and others. You can see that currently the opposite is happening e. g. AI slop companies pushing ads down to people. That's why we must make adblocking a basic human right - free access to information at all times.
It may sound strange right now, but people in the future will ask why we did not push back against these evil leeching companies before. Those fines are no longer enough nor sufficient - CEOs that break human rights need to go to jail for mandatory 5 years minimum, without any way to weasel out their way via money. Right now money is a way to get out - that system is broken. Anyone wondered how Epstein got so much money in the first place? Hint: he did not build it up by himself only.
cindyllm 13 hours ago [-]
[dead]
ricardobeat 12 hours ago [-]
Let’s not trivialize “human rights”, those are much more fundamental issues and this kind of discourse does not help achieve anything.
Where hash collisions matter is false positives, e.g. if hash(google.com) = hash(adserver.com). There does appear to be a web dashboard and /unblock api so should be straightforward to resolve.
This is a worrying trend on HN. Nearly for all submissions, there is one unquantified, unproven comment at the top saying it's AI, with just emotion and no proof. This is the new witch hunt, or karma farming without having to contribute anything positive to the discussion.
nearly all of the commits in the repo are co-authored by claude
If you have N entries where N >> 1, the probability of an arbitrary value colliding with the existing ones is Pcol(N+1) which is approx Pcol(N).
The math is correct, but it's the math for the question:
While the thing they probably want to check the collision math for is: As that's the risk you do the wrong action due to a hash collision. Importantly, the 2nd calculation is not going to be bounded by the size of your block list, it's going to be bounded by the number of domains overall.The former question isn't really useful - you can have a 100 entry blocklist with 100 collisions and it still works the same as if it had 0 collision since the action on all 100 is the same. That's more a problem for generating unique IDs per hash whereas this is flipped around because it's about classification.
If we assume uniform sampling for both the blocklist (size N) and the domain to be visited (not in the block list), the collision probability we care about is Pcol(N+1).
The number of existing domains does not matter, only the size of the block list.
Of course when we consider all 401 million domains there will be many more collisions, but in each case we only care about a collision in N+1
Edit: Or maybe they formulated that path to derivation and how to map it back but couldn't quickly find a tool able to approximate the birthday problem to that scale so they tried direct testing instead? If, e.g., they were slowly upping the values to see when a collision occurred it would make sense they ran into a collision earlier than would be expected as they (effectively) gave more than 1 trial. Or just random chance too if that's really the path they took I suppose :).
Edit2: After looking at the readme for far too long, I saw the original readme was in Japanese. Running that through a good translator actually clarifies or corrects a lot of the wording like saying "In the latter case, one other unlucky domain also gets blocked" rather than focusing on the number of domains over-blocked out of the 537000 only. With this version of the text I think you're definitely right, the math approach used should have given them the right number as the original text is already flipping things back to what happens with the one new query but they (apparently) just didn't do the actual calculation with it.
Sharp catch :).
It seems that the average number of expected collisions for 537k would be 0.13 and not 1:
the average value of pairs of individuals with the same birthday (or hash collisions) is E[X] = k*(k-1)/(2*n). Plugging in k=2^40 for 40 bits and n=537e3 I get their result of 0.1311
Thanks for the nerdsnipe, I ran out of my daily LLM mistake fixing quota !
[1] https://en.wikipedia.org/wiki/Birthday_problem#Average_numbe...
I read their statement as "there were 0 collisions when building a blocklist for 141k domains, and 1 collision for a blocklist of 537k domains", and the script points at that number coming directly from collisions within the blocklist.
1: https://github.com/M-Abozaid/esp32-c3-adblock/blob/238ef7b27...
This is a problem if the second one is actually a major useful domain. If HN and adserver have the same hash, you have a problem. This could be described as a false positive for HN.
> Where hash collisions matter is false positives, e.g. if hash(google.com) = hash(adserver.com)
This sounds like the same issue I describe above where you need to use Google but block adserver. Otherwise it's only a problem if you implemented allow lists, right? This is when you explicitly allow Google and implicitly allow adserver along with it because of the hash collision.
Also, PCB antennas are reasonably good for most indoor situations.
Great idea, and good for casual blocking, but I'm almost moving to an "allow list" mindset. This solution would probably work better for that, I wonder if the good parts of the internet would fit into an 140k list.
Since it's not ready, it'll be best if I point you to the sources I've aggregated from:
https://firebog.net/
https://github.com/hagezi/dns-blocklists
https://github.com/StevenBlack/hosts
https://github.com/jerryn70/GoodbyeAds
Link to my in-progress list file hosting: https://lists.uninvitedactivity.com/DNS/
00_Tiers directory:
I've aggregated the lists into four tiers: Base, Recommended, Aggressive, and Paranoid. And I've got two sets of these, one that includes Newly Registered Domains (Full directory), and one that doesn't include NRDs (NoNRD directory, because NRDs are ... heavy: Paranoid list with NRDs: 9.1 million records, without NRDs: 3.7 million records).
11_Allow directory:
A bunch of allow lists sourced from: https://discourse.pi-hole.net/t/commonly-whitelisted-domains.... I also use this list for an Allow list: https://github.com/anudeepND/whitelist
21_SpecificTopics directory:
Aggregation of lists under the topic that matches the file name. Fake News list needs to be updated as it contains a _way_ overbroad list that I need to remove from the aggregation.
For almost two decades I have been using an "allowlist", in the beginning it was via authoritative DNS. When proxy software began supporting mapping the DNS data into memory, I eliminates the need to use DNS
I have enough historical DNS data to assess own needs and it's much smaller than 140k
The popular blocklists can be quite strange if one takes the time to peruse them
They include what appear to be some very "bad parts of internet (www)" that I doubt most users would ever visit. There is more in these files than only common ad and tracking domains. As such one could call them "bloated"
It begs the question of how these blocklist maintainers even know about these bad parts
IMHO, the greatest exposure to ubiquitous ads, tracking and telemetry is through use of third party resolvers (in real-time, immediately preceding associated HTTP requests), where some process on the user's computer can potentially look up _any_ domain
Pretending that one can anticipate every possible undesired lookup, e.g., via a "blocklist", is IMHO a high risk approach
But it's popular. To each his own
Every user is different. Relying on a "blocklist" that attempts to anticipate every possible undesired domain is not for me. By contrast, the "allowlist" approach has served me reliably for many years. It sometimes suprises me how well it continues to work year after year
Philosophically, IMO, the contrast between "allowlist versus blocklist" is the difference between (a) "going after what one wants" from the www versus (b) "avoiding what one doesn't want"
The quantity of (a) can be relatively small and manageable. But (b) only grows larger over time, a tsunami full of garbage
I've always thought "ad blocking" was a misnomer. Instead of "blocking" what we as users want to do is "not send requests" to ad/tracking/telemetry servers
The www (HTTP) works via requests and responses. Generally, users make requests and publishers (including advertisers) receive the requests and (usually) send responses. Users are generally not receiving requests and few users would consciously send requests (including POSTing their private data) to ad, tracking or telemetry servers. It is "developers" who are making these requests for ads/tracking/telemetry, not users, by automatically generating them with the software they distribute to users (including Javascripts)
IMO, we don't "block" ads because we never requested them to begin with. What we do is ensure that the requests _made by developers_ do not succeed, we "block" requests that we never made
Segue into debate about user "agency"...
and every linux user should turn firewall on. and better download CISA hardening guides /best practices . or similar agency guides in countries other than US. they ARE free and mostly unversal, even if guide deals with other distro same principles apply to any system, just use AI to find right commands on your system.
99.999% of linux users who "converted to light of linux" based on nonsensical persuasion from perverts from LTT and other similar perverts, do not follow basic security practices. And these LTT perverts know this and they are laughing at your faces.
Windows in default configuration (as is installed without users turning off updates and other stuff off ) is 1000 milion times safer than linux for DESKTOP in default installed state. Which is state in which 99.9999% of loooosers sorry i meant to say users are using it in.. And im not even talking about totally irrational things to install like Crossover, Lutris and other nonsense based on WINE. I have 80 calls per month about infostealers, spyware in general which stem from using these nonsensical solutions full of old frameworsk (VB, .net etc ) and no security at all. Even official games, on EVERY platform, are scanning whole networks, combing thru your stuff (documents, pictures, webbrowsing activity... ) and noone cares, this is madness.
Distro maintainers should pull their heads from their asses too, how it is possibel to have firewall not enabled in OS in 2026?? HOOOOW. Not even talking about other stuff than like K12 stuff like firewall.
Depends on the user but not everyone is visiting new websites everyday
Even for those that are, the number of domain-IP mappings needed will be relatively small
Definitely under 140,000
Most DNS data I use is "static", it rarely changes. As such most times I don't have to make DNS queries. I store the domain-IP mappings in proxy memory; this is faster than DNS
No "blocklist" needed
If i would estimate it, I only take a look at 1/4 of the links where i read the comments.
Please just write the first sentence of your README yourself
I dislike when LLMs talk to me like that. I hate it when humans do it, confidently spewing their overconfident llm assumptions to others
I would only recommend using it instead of the standard Arduino Processing IDE.
Upstream dns for me is pretty quick. Google and cloudflare dns are ~5ms from me. But WiFi latency alone is ~8ms most of the time in my experience. On my fiber internet, pinging a dns server in some random upstream server miles away is lower latency than WiFi 10 feet away. But the real issue on WiFi is any packet loss at all adding 50-100ms to that at random depending on interference.
With DNS you are paying this latency cost all the time on nearly every request.
Surely not? macOS, Windows, and I think most of the big Linux distros, all cache DNS responses. Probably the web browser itself does too.
The real cost of using too much AI is nearly writing like them all the time.
The way Pi-Hole is used is to first determine whether a DNS should be blocked and only if it's not in that set, forward it to a remote resolver. I'd think the local resolver should be at least an order of magnitude less than the remote to justify it.
The developers could make users' lives easier by implementing a clustering option that syncs the 2 out of the box. For static deployments, where you don't change the config too often, it's still decent as it is.
I really do need to clean up my DNS setup a bit though. It's pretty wacky right now.
Store the domains as sorted 40-bit hashes in flash and binary-search them.
140,000+ domains fit in ~0.7 MB of flash..."
Brilliant! (Although, what about potential hash collisions with domain names that shouldn't be blocked? Those are probably few and far between, so the utility of the block overall far outweighs any few false positives that might arise -- so I reiterate my claim of "brilliant"!)
Also (if you're a compression nerd like I am), it may be possible to analyze this table of 40-bit hashes further to find the longest binary substring that repeats the most across all of them (or binary substrings that just repeat a whole lot), and design a Huffman-like (or other tree or tree-like) data structure around those, and possibly compress the data further...
Of course, then the simplicity and elegance of the Binary Search might be lost, but it might be an interesting optimization exercise to know just how far that data could be compressed... "Hey Claude, try to optimize the compression of this data using alternative compression data structures, and report back to me!" (Or something like that! You know, get the AI coding agents to try different things -- or attempt to code it on your own (always better for understanding!), etc., etc.!)
Anyway, brilliant!
after all, why use opnSense; just use suricata and $PROPER_X! why use TrueNAS; just use $DISTRO and ZFS!
Jokes aside, very impressive that it works!
It may sound strange right now, but people in the future will ask why we did not push back against these evil leeching companies before. Those fines are no longer enough nor sufficient - CEOs that break human rights need to go to jail for mandatory 5 years minimum, without any way to weasel out their way via money. Right now money is a way to get out - that system is broken. Anyone wondered how Epstein got so much money in the first place? Hint: he did not build it up by himself only.