Merge branch 'master' of https://github.com/linuxlefty/hosts into linuxlefty-master

# By Peter Naudus
# Via Peter Naudus
* 'master' of https://github.com/linuxlefty/hosts:
  Normalizing hosts and better duplicate detction added
  Added hpHosts (hosts-file.net) as a source
  Added support for ZIP compressed files
  Removing vim swap file
  Added yoyo.org adserver lists as data source
  Refreshed hosts from malwaredomainlist.com, mvps.org, and someonewhocares.org

Conflicts:
	data/malwaredomainlist.com/hosts
	data/mvps.org/hosts
	data/someonewhocares.org/hosts
	hosts
	readme.md
This commit is contained in:
Steven Black
2014-05-31 00:38:11 -04:00
7 changed files with 449311 additions and 6 deletions
File diff suppressed because it is too large Load Diff
+1
View File
@@ -0,0 +1 @@
http://hosts-file.net/download/hosts.zip
+2503
View File
File diff suppressed because it is too large Load Diff
+1
View File
@@ -0,0 +1 @@
http://pgl.yoyo.org/adservers/serverlist.php?hostformat=hosts&mimetype=plaintext
+5 -2
View File
@@ -2,7 +2,7 @@
This repo consolidates several reputable `hosts` files and consolidates them into a single hosts file that you can use.
**Currently this hosts file contains 25136 unique entries.**
**Currently this hosts file contains 465030 unique entries.**
## Source of host data amalgamated here
@@ -10,6 +10,9 @@ Currently the `hosts` files from the following locations are amalgamated:
* MVPs.org Hosts file at [http://winhelp2002.mvps.org/hosts.htm](http://winhelp2002.mvps.org/hosts.htm), updated monthly, or thereabouts.
* Dan Pollock at [http://someonewhocares.org/hosts/](http://someonewhocares.org/hosts/) updated regularly.
* Malware Domain List at [http://www.malwaredomainlist.com/](http://www.malwaredomainlist.com/), updated regularly.
* Peter Lowe at [http://pgl.yoyo.org/adservers/](http://pgl.yoyo.org/adservers/), updated regularly.
* hpHosts at [http://hosts-file.net/](http://hosts-file.net/), updated regularly
* My own small list in raw form [here](https://raw.github.com/StevenBlack/hosts/master/data/StevenBlack/hosts).
You can add any additional sources you'd like under the data/ directory. Provide a copy of the current `hosts` file and a file called
@@ -74,4 +77,4 @@ and run:
### Linux
Open a Terminal and run:
`/etc/rc.d/init.d/nscd restart`
`/etc/rc.d/init.d/nscd restart`
+3
View File
@@ -10,6 +10,9 @@ Currently the `hosts` files from the following locations are amalgamated:
* MVPs.org Hosts file at [http://winhelp2002.mvps.org/hosts.htm](http://winhelp2002.mvps.org/hosts.htm), updated monthly, or thereabouts.
* Dan Pollock at [http://someonewhocares.org/hosts/](http://someonewhocares.org/hosts/) updated regularly.
* Malware Domain List at [http://www.malwaredomainlist.com/](http://www.malwaredomainlist.com/), updated regularly.
* Peter Lowe at [http://pgl.yoyo.org/adservers/](http://pgl.yoyo.org/adservers/), updated regularly.
* hpHosts at [http://hosts-file.net/](http://hosts-file.net/), updated regularly
* My own small list in raw form [here](https://raw.github.com/StevenBlack/hosts/master/data/StevenBlack/hosts).
You can add any additional sources you'd like under the data/ directory. Provide a copy of the current `hosts` file and a file called
+28 -4
View File
@@ -14,6 +14,8 @@ import subprocess
import sys
import tempfile
import urllib2
import zipfile
import StringIO
# Project Settings
BASEDIR_PATH = os.path.dirname(os.path.realpath(__file__))
@@ -23,6 +25,7 @@ UPDATE_URL_FILENAME = 'update.info'
SOURCES = os.listdir(DATA_PATH)
README_TEMPLATE = BASEDIR_PATH + '/readme_template.md'
README_FILE = BASEDIR_PATH + '/readme.md'
TARGET_HOST = '0.0.0.0'
# Exclusions
EXCLUSION_PATTERN = '([a-zA-Z\d-]+\.){0,}' #append domain the end
@@ -115,7 +118,16 @@ def updateAllSources():
continue;
print 'Updating source ' + source + ' from ' + updateURL
updatedFile = urllib2.urlopen(updateURL)
updatedFile = updatedFile.read()
if '.zip' in updateURL:
updatedZippedFile = zipfile.ZipFile(StringIO.StringIO(updatedFile))
for name in updatedZippedFile.namelist():
if name in ('hosts', 'hosts.txt'):
updatedFile = updatedZippedFile.open(name).read()
break
updatedFile = string.replace( updatedFile, '\r', '' ) #get rid of carriage-return symbols
dataFile = open(DATA_PATH + '/' + source + '/' + DATA_FILENAMES, 'w')
@@ -151,7 +163,7 @@ def removeDups(mergeFile):
finalFile = open(BASEDIR_PATH + '/hosts', 'w+b')
mergeFile.seek(0) # reset file pointer
rules_seen = set()
hostnames = set()
for line in mergeFile.readlines():
if line[0].startswith("#") or line[0] == '\n':
finalFile.write(line) #maintain the comments for readability
@@ -159,15 +171,27 @@ def removeDups(mergeFile):
strippedRule = stripRule(line) #strip comments
if matchesExclusions(strippedRule):
continue
if strippedRule not in rules_seen:
finalFile.write(line)
rules_seen.add(strippedRule)
hostname, normalizedRule = normalizeRule(strippedRule) # normalize rule
if normalizedRule and hostname not in hostnames:
finalFile.write(normalizedRule)
hostnames.add(hostname)
numberOfRules += 1
else:
finalFile.write(line)
mergeFile.close()
return finalFile
def normalizeRule(rule):
result = re.search(r'^\s*(\d+\.\d+\.\d+\.\d+)\s+([\w\.-]+)(.*)',rule)
if result:
target, hostname, suffix = result.groups()
return hostname, "%s\t%s%s\n" % (TARGET_HOST, hostname, suffix)
print '==>%s<==' % rule
return None, None
def finalizeFile(finalFile):
writeOpeningHeader(finalFile)
finalFile.close()