This is probably the densest paper we've read. It's an in-depth analysis of TCP's behavior. It attempts to analyze the assumptions we have when building TCP, just as the network's infrequent reordering to the packet error rate. It's an immense work.
On a note, I'm going to start a list of worst project/protocol names. New on the list: SACK. SACKs are not new to me, but this is a new list and I've recently been reminded of it.
The usual complaints apply to this work. It's old, we're not sure how relevant the findings are to the modern internet. There were many reorderings and bursty losses due to routers updating their routes. This is presumably less likely now. Follow-on work would be great to read, as it would show what changes were implemented in reaction to these finding (not many I'd wager).
I really like these deep, measurement papers. There's so much nonsense about building a system when there's such a light understanding of the problem you're trying to solve. This work likely led to tens of future projects for exactly that reason. There are plenty of engineers in systems, we need more analysis.
Particularly, I want more social science techniques in systems and networking. This is probably a deeper analysis of user cases. We build so many technologies that are solving problems that don't exist. They get into SIGCOMM, put in a binder and on a CV, and that's it. Though, to be fair, that may just be academic problem...
Anyhow, I feel like I had something else of value to say about this work, but I can't remember. Perhaps I'll update once I remember.
Wednesday, November 12, 2008
Thursday, November 6, 2008
i3
Matei lied to me and told me that this was the better paper of the two.
i3 is a fairly obvious scheme to provide mobility and multicast with an overlay. There's a few interesting implementation details that they tweezed out of the idea, but none are crucial for the actual design.
Again I come to question that has haunted me through this class, adoption. Overlays are nice for adoption, as you presumably control each client. This works particularly well for p2p, less so for others.
A big part of this paper was being general and applying to traditional servers, which don't need this. You only pay a performance penalty if there is no mobility or multicast. The lack of these is the common case. I also think it will remain the common case as we deploy more wide area wireless networks.
So, this really fits mobility and multicast in p2p networks. Each has interesting use cases I think, multicast is nice to steal stuff and mobility will allow a large class of applications on personal communication devices.
i3 is a fairly obvious scheme to provide mobility and multicast with an overlay. There's a few interesting implementation details that they tweezed out of the idea, but none are crucial for the actual design.
Again I come to question that has haunted me through this class, adoption. Overlays are nice for adoption, as you presumably control each client. This works particularly well for p2p, less so for others.
A big part of this paper was being general and applying to traditional servers, which don't need this. You only pay a performance penalty if there is no mobility or multicast. The lack of these is the common case. I also think it will remain the common case as we deploy more wide area wireless networks.
So, this really fits mobility and multicast in p2p networks. Each has interesting use cases I think, multicast is nice to steal stuff and mobility will allow a large class of applications on personal communication devices.
Wednesday, November 5, 2008
DOA
Woah. First thing, don't name a project DOA. Name it something better, like PANTS.
Secondly, this is 16 pages about a fairly simple idea. The diea is to generalize middleboxes in some way, dealing with all of the issues the things bring up. The most common issue is addressing machines inside of NATs, as their addresses are meaningless outside of the LAN. These guys define a new namespace that is global and flat, fudge a DNS scheme, and allow the packets to give a source routing like scheme.
This is an interesting point though. Essentially, source routing solves all of this, right? I mean, if I send a source routed packet to a NAT box, with (192.168.0.10) as the last source to route to, it would work. Assuming local DNS (such as avahi) you could make the last source be: (darth-vader) or (waterloo) and get the packet there without global naming. I suppose that the NAT will need to be addressed in some meaningful manner, but those usually are globally unique.
Huh, that might be a good idea.
Anyhow, this work was good, but it's got that "never going to be used" problem, as it requires both sides use it to get an advantage. I don't think this is inherent to the solution either, as my thing above is a slightly more reasonable version of this. For instance, a VoIP app may just broadcast a "register" packet saying "hey! I'm listening on this port for VoIP, if you get one, give it here". Then you've done some of the delegation they speak of, without the new naming scheme. You can get looked up via traditional DNS and bam, way around that punching a hole crap.
Lastly, note that skype fixed this with a big server in the real world. Both clients talk to it, and then skype tells each client what port the other has open. Doesn't work for pure P2P, unless you can negotiate a P2P node to do this for you...
Secondly, this is 16 pages about a fairly simple idea. The diea is to generalize middleboxes in some way, dealing with all of the issues the things bring up. The most common issue is addressing machines inside of NATs, as their addresses are meaningless outside of the LAN. These guys define a new namespace that is global and flat, fudge a DNS scheme, and allow the packets to give a source routing like scheme.
This is an interesting point though. Essentially, source routing solves all of this, right? I mean, if I send a source routed packet to a NAT box, with (192.168.0.10) as the last source to route to, it would work. Assuming local DNS (such as avahi) you could make the last source be: (darth-vader) or (waterloo) and get the packet there without global naming. I suppose that the NAT will need to be addressed in some meaningful manner, but those usually are globally unique.
Huh, that might be a good idea.
Anyhow, this work was good, but it's got that "never going to be used" problem, as it requires both sides use it to get an advantage. I don't think this is inherent to the solution either, as my thing above is a slightly more reasonable version of this. For instance, a VoIP app may just broadcast a "register" packet saying "hey! I'm listening on this port for VoIP, if you get one, give it here". Then you've done some of the delegation they speak of, without the new naming scheme. You can get looked up via traditional DNS and bam, way around that punching a hole crap.
Lastly, note that skype fixed this with a big server in the real world. Both clients talk to it, and then skype tells each client what port the other has open. Doesn't work for pure P2P, unless you can negotiate a P2P node to do this for you...
Sunday, November 2, 2008
DNS Caching
This was a dense work from MIT and KAIST where they gathered traces of DNS traffic to determine the effectiveness of DNS caching. There were two primary findings: DNS requests follow the known zipfian distribution, and low TTL DNS entries do not really affect the caching.
This was really really dense with a whole lot of figures and percentages. I think I followed most of it, with the argument that the zipfian distribution is correct being the most obvious. They argue it shows that caching is not that important, as most users only use the cache for a few minutes as they browse the same site. This is important as servers are using TTL-based load balancing.
Note that this only applies for addressing, not forwarding. I think this is a crucial point, as it allows us to distribute the load at that level. Since it's likely amazon.com owns that hardware, TTL multiplexing makes a lot of sense.
Interesting bit about scaling out DNS. I'm always interested in how these things have changed, as I feel like the internet has become much more stable since 2000. Then, the obvious question is: how did they stabilize it?
This was really really dense with a whole lot of figures and percentages. I think I followed most of it, with the argument that the zipfian distribution is correct being the most obvious. They argue it shows that caching is not that important, as most users only use the cache for a few minutes as they browse the same site. This is important as servers are using TTL-based load balancing.
Note that this only applies for addressing, not forwarding. I think this is a crucial point, as it allows us to distribute the load at that level. Since it's likely amazon.com owns that hardware, TTL multiplexing makes a lot of sense.
Interesting bit about scaling out DNS. I'm always interested in how these things have changed, as I feel like the internet has become much more stable since 2000. Then, the obvious question is: how did they stabilize it?
DNS
This paper was strange. I know DNS, and this paper actually made me a little more confused. I expect it was mostly terminology. Zones, for instance, are a strange name for administrative domains.
The paper itself was great, as it dealt heavily with an issue very close to my heart; technology adoption. As I work in Information and Communication Technology for Development (ICTD) getting users of my technology is of paramount importance.
The HOSTS.TXT file was a perfect example of "good enough" technology, which is the biggest obstacle to adoption. It's just not worth it to the users unless the benefits of the new technology outweigh the losses from adoption. This is why it's nearly impossible to get new networking technology adopted. The advantages are too small compared to the risks.
The DNS folks lamented a great deal over this fact. The users who did not switch caused tons of problems, and the users who only partially transitioned caused the most. I think that Berkeley TCP implementations were probably the definitive reason that TCP is so widely adopted.
I'm not sure exactly what I'm trying to say. I just want a more involved discussion on how to push network technology. Right now you have to implement it at user level, that way people don't need to monkey with their systems to use it.
I'm rambling. I'l bring this up in class.
The paper itself was great, as it dealt heavily with an issue very close to my heart; technology adoption. As I work in Information and Communication Technology for Development (ICTD) getting users of my technology is of paramount importance.
The HOSTS.TXT file was a perfect example of "good enough" technology, which is the biggest obstacle to adoption. It's just not worth it to the users unless the benefits of the new technology outweigh the losses from adoption. This is why it's nearly impossible to get new networking technology adopted. The advantages are too small compared to the risks.
The DNS folks lamented a great deal over this fact. The users who did not switch caused tons of problems, and the users who only partially transitioned caused the most. I think that Berkeley TCP implementations were probably the definitive reason that TCP is so widely adopted.
I'm not sure exactly what I'm trying to say. I just want a more involved discussion on how to push network technology. Right now you have to implement it at user level, that way people don't need to monkey with their systems to use it.
I'm rambling. I'l bring this up in class.
Thursday, October 30, 2008
DHTs
I've decided to put both DHT papers into the same post, as they covered a lot of the same material.
These were papers on distributed hash tables. They what they seem to be, but there are some clever tricks in each individual one. You hash an address and this then maps to another machine via a series of lookups.
What I don't understand is how these things help P2P, as is argued at such depth. Since you don't get to organize the data in a pure P2P network, you can't really create any address -> data mappings. Both are handed to you, you can only create a search service that gives you one if given the other.
I do like DHTs as DNS replacements though, that seems like the proper model of usage. This is an actual problem too, as I can remember comcast's DNS server choking and access to goggle being hindered.
I guess I don't have much to write as we read the Chord paper for 262. All of my clever insight is gone.
These were papers on distributed hash tables. They what they seem to be, but there are some clever tricks in each individual one. You hash an address and this then maps to another machine via a series of lookups.
What I don't understand is how these things help P2P, as is argued at such depth. Since you don't get to organize the data in a pure P2P network, you can't really create any address -> data mappings. Both are handed to you, you can only create a search service that gives you one if given the other.
I do like DHTs as DNS replacements though, that seems like the proper model of usage. This is an actual problem too, as I can remember comcast's DNS server choking and access to goggle being hindered.
I guess I don't have much to write as we read the Chord paper for 262. All of my clever insight is gone.
Monday, October 27, 2008
Active Networks
I still don't really know what active networks are. Wiki says:
"The active network architecture is composed of execution environments (similar to a unix shell that can execute active packets), a node operating system capable of supporting one or more execution environments. It also consists of active hardware, capable of routing or switching as well as executing code within active packets. This differs from the traditional network architecture which seeks robustness and stability by attempting to remove complexity and the ability to change its fundamental operation from underlying network components. Network processors are one means of implementing active networking concepts. Active networks have also been implemented as overlay networks."
That didn't help too much. I suppose it's reasonable that the applications should be allowed to reform the network. However, as my previous post said, this isn't going to optimize throughput, and thus won't be of huge value.
Is this a formative work for overlay networks?
Rambling aside, the paper itself was a survey of the lessons learned in implementing such a system. Not terribly valuable, in my own view.
Lastly, I am intrigued by the idea, at least for the DTN case. DTN's are highly mobile, with the network shifting around all the time. It may be valuable to modify the network architecture in this case. Then again, I barely understand this to begin with. I'll muse more after lecture.
"The active network architecture is composed of execution environments (similar to a unix shell that can execute active packets), a node operating system capable of supporting one or more execution environments. It also consists of active hardware, capable of routing or switching as well as executing code within active packets. This differs from the traditional network architecture which seeks robustness and stability by attempting to remove complexity and the ability to change its fundamental operation from underlying network components. Network processors are one means of implementing active networking concepts. Active networks have also been implemented as overlay networks."
That didn't help too much. I suppose it's reasonable that the applications should be allowed to reform the network. However, as my previous post said, this isn't going to optimize throughput, and thus won't be of huge value.
Is this a formative work for overlay networks?
Rambling aside, the paper itself was a survey of the lessons learned in implementing such a system. Not terribly valuable, in my own view.
Lastly, I am intrigued by the idea, at least for the DTN case. DTN's are highly mobile, with the network shifting around all the time. It may be valuable to modify the network architecture in this case. Then again, I barely understand this to begin with. I'll muse more after lecture.
Subscribe to:
Posts (Atom)